RabbitMQ, NATS and Kafka queues: revisit your message queue before you add another broker

A health insurance company in Hyderabad processes claims through RabbitMQ: three nodes on 3.13 with classic mirrored queues, about 40 lakh messages a day, from claim documents to OCR, OCR results to fraud checks and approvals to payments. The upgrade to 4.x has been postponed twice because “mirroring goes away”. The field team now wants NATS for hospital kiosks, and the data platform team, which runs Kafka for analytics, says Kafka now has queues, so why not move everything there?

That is a fair question, and the honest answer is not “pick the fastest”. A queue is a contract about what happens when a consumer fails: when a message comes back, how many times, in what order, and where it goes when it can never succeed. The three brokers answer differently, and two changed their answers in the last eighteen months. This post compares RabbitMQ, NATS JetStream and Kafka share groups on queue semantics, replication, operations, security and licences, with short notes on the alternatives. Facts come from release notes, KIPs, official docs, licence files and CVE records, checked on 8 October 2026; the drill is my own run.

The short version

  • RabbitMQ 4.3 removed Mnesia: Khepri, a Raft-based store, is now the only metadata store, so a cluster needs a majority of nodes online. You can reach 4.3 only from 4.2.x, and classic queue mirroring has been gone since 4.0.
  • RabbitMQ 4.3 quorum queues count only failed deliveries against the delivery limit. In my drill, basic.nack with requeue redelivered one message 30 times and never dead-lettered it, while basic.reject sent it to the dead letter queue after the fourth delivery.
  • NATS stays an Apache 2.0 project in CNCF. In 2025 Synadia announced a plan to move the server to the Business Source License; on 1 May 2025 it instead agreed to assign the NATS trademarks to the Linux Foundation.
  • JetStream has no built-in dead letter queue. When MaxDeliver is reached it publishes an advisory and leaves the message in the stream, so you build the dead letter path yourself.
  • Kafka share groups (KIP-932) became production-ready in Kafka 4.2.0 (February 2026). There is still no dead letter topic in 4.3.1, and a new share group starts at the latest offset by default.
  • RabbitMQ had more than 50 CVE records published in September 2026 alone. Patch planning is part of the choice, not an afterthought.

Where things stand on 8 October 2026

RabbitMQ, NATS, Pulsar, LavinMQ, ActiveMQ and pgmq dates are from GitHub release pages in IST; Kafka dates are from the Apache Kafka release announcements.

Project

Latest release

What it is

Licence

RabbitMQ

4.3.6 on 14 September 2026

Erlang broker: AMQP 0-9-1, AMQP 1.0, MQTT, STOMP, streams

MPL 2.0 (server and core plugins)

NATS server

2.15.0 on 17 September 2026; 2.14.7 on 15 September

Go server with JetStream persistence

Apache 2.0

Apache Kafka

4.3.1 (25 June 2026); 4.2.2 (29 September 2026)

Partitioned log, now with share groups

Apache 2.0

Apache Pulsar

5.0.0, 4.2.5 and 4.0.14 on 6 October 2026

Brokers on BookKeeper, queue and stream subscriptions

Apache 2.0

Apache Artemis (formerly ActiveMQ Artemis)

2.57.0, tagged 9 September 2026

Java multi-protocol broker

Apache 2.0

Apache ActiveMQ Classic

6.3.2 (29 August) and 6.2.10 (22 September 2026)

The original Java broker

Apache 2.0

LavinMQ

2.10.2 on 7 October 2026

AMQP 0-9-1 and MQTT broker in Crystal

Apache 2.0

pgmq

1.13.0 on 7 September 2026

PostgreSQL extension for SQS-style queues, PostgreSQL 14 to 18

PostgreSQL License

Kafka 4.4.0 was at its fourth release candidate on 6 October 2026, and NATS 2.15.1 at its second.

Date

Event

15 March 2018

NATS accepted into CNCF at incubating level

5 June 2019

NATS server 2.0.0: the repository and binary gnatsd become nats-server

1 October 2019

RabbitMQ 3.8.0 introduces quorum queues

15 March 2021

NATS server 2.2.0 introduces JetStream

23 July 2021

RabbitMQ 3.9.0 introduces streams

23 February 2024

RabbitMQ 3.13.0 offers Khepri as an alternative to Mnesia

19 September 2024

RabbitMQ 4.0.1: classic queue mirroring removed, AMQP 1.0 a core protocol, Khepri fully supported

18 March 2025

Kafka 4.0.0 ships Queues for Kafka (KIP-932) as early access

April to 1 May 2025

CNCF and Synadia dispute over NATS, ending in a trademark agreement

4 September 2025

Kafka 4.1.0 moves share groups to preview

27 October 2025

RabbitMQ 4.2.0 makes Khepri the default for new clusters

17 February 2026

Kafka 4.2.0 declares share groups production-ready

23 April 2026

RabbitMQ 4.3.0 removes Mnesia; quorum queues gain delayed retry, consumer timeout and strict priorities

31 July 2026

RabbitMQ 4.2 community support ends

17 September 2026

NATS server 2.15.0 with a desired-state metalayer and evacuation endpoints

Queue or log: why this question keeps coming back

A log keeps records in order and lets each consumer group move its own offset. A queue hands each message to one worker, waits for a verdict, and redelivers on failure. Each broker has now moved into the others’ territory: RabbitMQ has streams, NATS has JetStream, and Kafka has share groups. The earlier post on Kafka and Redpanda covered Kafka as a streaming log; this one is about the queue side.

The Hyderabad workload is queue-shaped. Each claim document is an independent job; some take 2 seconds, some 4 minutes; some fail while the OCR service is down, some fail forever because the PDF is corrupt. What matters is per-message acknowledgement, a sane retry schedule, a dead letter path, and no ordering promise you do not need.

Queue semantics, side by side

Question

RabbitMQ 4.3 quorum queue

NATS JetStream consumer

Kafka 4.3 share group

PostgreSQL table or pgmq

How a worker answers

ack, nack, reject; AMQP 1.0 accepted, released, rejected, modified

ack, nak (optionally with delay), in-progress, term

ACCEPT, RELEASE, REJECT, RENEW

Delete or archive the row

Redelivery when the worker goes silent

On connection loss, or after a consumer timeout (30 minutes by default)

After AckWait, 30 s by default

After the acquisition lock, 30 s by default

pgmq: when the visibility timeout ends

Retry limit

delivery-limit, default 20 since 4.0

MaxDeliver, default -1 (unlimited)

Delivery count limit, default 5 (2 to 10)

Your code, using read_ct

Where poison messages go

Dead letter exchange; at-least-once mode available

Nowhere: an advisory is published, the message stays

Archived; no dead letter topic yet

Your archive or error table

Back-off between retries

Delayed retry since 4.3, linear

Consumer BackOff list, or nak with delay

None built in

Your code

Ordering

FIFO per queue, but redeliveries break it

Stream order; redeliveries arrive after later messages unless MaxAckPending is 1

Increasing offsets within a batch only

Whatever your ORDER BY gives

Priority

32 strict levels on quorum queues since 4.3

Priority groups for consumers, not messages

None

An index on a priority column

Per-message TTL

Yes

Nats-TTL header since 2.11

No; KIP-932 mentions it only as a possible extension

Your code

Two things stand out. First, every option is at-least-once. A worker that charges a card or sends an SMS must be safe to run twice, and the post on idempotency keys covers how to make it so. Second, the defaults differ widely. Moving from RabbitMQ to JetStream changes your retry limit from 20 to unlimited unless you set it; moving to Kafka share groups changes it to 5 and removes the dead letter queue.

RabbitMQ 4.x: the broker you knew has changed

Mirroring is gone; quorum queues and streams replicate

RabbitMQ 4.0 removed classic queue mirroring after three years of deprecation. Classic queues still work, but with one replica. For replicated data you now choose quorum queues, which replicate with Raft, or streams, which are replicated append-only logs. A quorum queue has up to three members by default. The docs promise that a message confirmed to the publisher is not lost while a majority of the nodes hosting the queue survive, and promise nothing for unconfirmed messages. Publisher confirms are not optional for claims data.

Khepri is the only metadata store

Users, vhosts, queues, bindings and policies live in the metadata store. 4.2.0 made Khepri the default for new clusters; 4.3.0 removed Mnesia and the old partition handling settings. The 4.3 notes spell out the change: a majority of nodes must be online for the cluster to be available. Two-node clusters were always a bad idea; now they are an outage waiting to happen.

AMQP 1.0 is native, and 4.3 tightens old habits

Since 4.0, AMQP 1.0 is a core protocol rather than a plugin, and 4.1 and 4.2 added filter expressions for streams. 4.3 denies some deprecated features by default (transient non-exclusive queues, global QoS) and removes CQv1 storage. Test your client libraries against 4.3 before the upgrade window, not during it.

What 4.3 changed for quorum queues

The 4.3.0 release notes list delayed retry with linear back-off (min_delay x delivery_count, capped at max_delay), consumer timeouts handled inside the queue, strict priorities with 32 levels, and a compact message reference format that the RabbitMQ team says halves per-message memory overhead in many scenarios. That last one is the vendor’s own claim.

The change that can surprise you is subtle. Since 4.3, the delivery limit is based on delivery-count, which increases only on failures: basic.reject, an AMQP 1.0 rejected or modified with delivery-failed=true, or a client crash. A basic.nack with requeue, an AMQP 1.0 released, and even a consumer timeout increase only x-acquired-count. My drill shows the result: a consumer that nacks a poison message loops on it forever, even with a delivery limit of 3. If your consumers use nack for permanent failures, switch to reject (or nack without requeue) and keep a dead letter target.

A policy keeps this out of application code. I tested this on 4.3.6: it produced six deliveries with gaps of 1, 2, 3, 4 and 5 seconds, then moved the message to dlq.orders.

rabbitmqctl set_policy orders-qq '^orders\.' \
  '{"delivery-limit": 5,
    "dead-letter-exchange": "",
    "dead-letter-routing-key": "dlq.orders",
    "dead-letter-strategy": "at-least-once",
    "overflow": "reject-publish",
    "delayed-retry-type": "failed",
    "delayed-retry-min": 1000,
    "delayed-retry-max": 30000}' \
  --apply-to quorum_queues

Keep the dead letter queue’s name outside the policy pattern. The at-least-once strategy needs overflow set to reject-publish; the default at-most-once can lose dead-lettered messages.

Support policy and ownership

The rabbitmq.com footer carries a Broadcom copyright, and Broadcom sells commercial support. The community support policy says patches for older series are not made available to non-paying users, with a possible exception for very high severity CVEs. On 8 October 2026 the release information page lists 4.2.10 (17 August), but GitHub’s latest public 4.2 release is 4.2.9 (20 July). Without a contract, “stay on 4.2 for a while” means “stay unpatched”. Plan to follow 4.3.x.

NATS and JetStream

From gnatsd to a CNCF project

NATS began as a lightweight pub/sub server; CNCF’s project page dates the first commit to 30 October 2010. The server was called gnatsd until 2.0.0 (June 2019). NATS joined CNCF in March 2018 and is still incubating. JetStream, the persistence layer, arrived in 2.2.0 in March 2021. Core NATS without JetStream is at-most-once: if no subscriber is listening, the message is gone.

In 2025 the project’s governance was briefly in doubt. CNCF wrote that Synadia, the company behind most NATS development, had told it of a plan to withdraw NATS from the foundation and adopt the Business Source License for the server. On 1 May 2025, CNCF and Synadia announced an agreement: Synadia assigns its NATS trademark registrations to the Linux Foundation, CNCF keeps the nats.io domain and GitHub repositories, and development continues under Apache 2.0. The LICENSE file in nats-server is still Apache 2.0.

How JetStream behaves as a queue

A stream stores messages; a consumer is a view with its own delivery state. For queue work you use a pull consumer with explicit acks, often on a stream with WorkQueue retention, which deletes a message once it is acknowledged. The defaults from the docs: AckWait 30 seconds, MaxDeliver -1 (no limit), MaxAckPending 1000. A BackOff list replaces AckWait for timed-out deliveries, and a plain nak redelivers immediately unless the client attaches a delay. term tells the server never to redeliver.

The docs are direct: JetStream has no built-in dead letter queue. When a message hits MaxDeliver, the server stops delivering it and publishes an advisory on $JS.EVENT.ADVISORY.CONSUMER.MAX_DELIVERIES.<stream>.<consumer>. In my drill, the message stayed in the WorkQueue stream after the advisory, still using storage, and the advisory arrived only while a worker had a pull request pending. Here is the pattern I tested on 2.15.0 with nats-py 2.16.0; it ended with one message in DLQ and none in ORDERS:

import asyncio, json, nats
from nats.js.api import StreamConfig, ConsumerConfig, RetentionPolicy, AckPolicy

async def main():
    nc = await nats.connect("nats://127.0.0.1:4222")
    js = nc.jetstream()
    await js.add_stream(StreamConfig(name="ORDERS", subjects=["orders.>"],
                                     retention=RetentionPolicy.WORK_QUEUE))
    await js.add_stream(StreamConfig(name="DLQ", subjects=["dlq.>"]))

    async def to_dlq(msg):  # JetStream has no built-in DLQ: build one
        adv = json.loads(msg.data)
        raw = await js.get_msg("ORDERS", adv["stream_seq"])
        await js.publish("dlq.orders", raw.data)
        await js.delete_msg("ORDERS", adv["stream_seq"])
    await nc.subscribe("$JS.EVENT.ADVISORY.CONSUMER.MAX_DELIVERIES.ORDERS.*", cb=to_dlq)

    await js.add_consumer("ORDERS", ConsumerConfig(
        durable_name="billing", ack_policy=AckPolicy.EXPLICIT,
        max_deliver=4, backoff=[1, 2, 4], filter_subject="orders.bill"))
    await js.publish("orders.bill", b"invoice-77")
    sub = await js.pull_subscribe_bind("billing", "ORDERS")
    for _ in range(8):                 # a worker that keeps polling, never acks
        try:
            await sub.fetch(1, timeout=2)
        except nats.errors.TimeoutError:
            pass
    print((await js.stream_info("DLQ")).state.messages,
          (await js.stream_info("ORDERS")).state.messages)
    await nc.close()

asyncio.run(main())

A core subscription misses advisories published while the handler is down, so in production capture advisories in a stream and read them with a durable consumer.

What changed from 2.11 to 2.15

2.11.0 (March 2025) added per-message TTLs, consumer priority groups, consumer pausing and ingest rate limiting. 2.12.0 (September 2025) added atomic batches, counters and delayed message scheduling. 2.14.0 (April 2026; 2.13 was skipped) added cron-style schedules and a consumer reset API. 2.15.0 added a desired-state metalayer for safer stream moves, evacuation endpoints and a default limit of 1,000 consumers per stream; applications that create many consumers should set max_consumers explicitly.

Kafka share groups: queues, the Kafka way

KIP-932 does not add a queue object. It adds share groups: consumers read the same partitions cooperatively, each record is locked to one consumer for a time, and every record gets its own acknowledgement and delivery count. Parallelism is no longer capped by the partition count.

It took three releases: early access in 4.0.0 (March 2025, disabled by default and not for production), preview in 4.1.0 (September 2025), and production-ready in 4.2.0 (February 2026), which also added the RENEW acknowledgement for long jobs and share-partition lag metrics. 4.3.0 (May 2026) added group-level settings such as share.delivery.count.limit and share.renew.acknowledge.enable.

The defaults from the KIP: a 30-second acquisition lock (1 second to 1 hour), a delivery count limit of 5 (between 2 and 10), and at most 2,000 in-flight records per share partition. After the limit, a record is archived, not dead-lettered. KIP-1191, which adds dead letter topics for share groups, is accepted but tied to share.version=2, and my 4.3.1 broker supports only version 1. Ordering is weak by design: records within one batch arrive in offset order, nothing more. Transactional (exactly-once) acknowledgement is listed as future work.

Three practical points. A fresh 4.3.1 cluster had share.version=1 without my setting it; on an upgraded cluster, check kafka-features.sh describe. With fewer than three brokers, set share.coordinator.state.topic.replication.factor and share.coordinator.state.topic.min.isr to 1 for the internal __share_group_state topic. And a new share group starts at the latest offset, so in my drill it ignored 2,000 existing records until I changed the group setting:

kafka-configs.sh --bootstrap-server 127.0.0.1:19092 --alter \
  --entity-type groups --entity-name g-release \
  --add-config share.auto.offset.reset=earliest
kafka-console-share-consumer.sh --bootstrap-server 127.0.0.1:19092 \
  --topic jobs --group g-release --release \
  --formatter-property print.delivery=true

With --release, the consumer handed the same record back five times (Delivery:1 to Delivery:5) and then never saw it again.

Share groups suit teams already on Kafka with independent jobs; they fit less well if you need dead letters, TTLs, priorities or delays today.

The others, briefly

  • Apache Pulsar has shared subscriptions, negative acks, redelivery back-off and a dead letter policy. Its docs warn that the negative-ack redelivery counter lives only in memory and resets on broker restart, topic unload or consumer disconnect; they recommend reconsumeLater with retry enabled for a reliable limit.
  • Apache Artemis and ActiveMQ Classic. ActiveMQ Artemis is now called Apache Artemis. Both are mature JMS brokers for Java estates, and both had serious 2026 CVEs, below.
  • LavinMQ, from CloudAMQP, speaks AMQP 0-9-1 and MQTT 3.1.x and is written in Crystal. Check the features your clients use before treating it as a RabbitMQ replacement.
  • Redis and Valkey streams offer consumer groups, a pending entries list and XAUTOCLAIM (since 6.2.0) to take over idle messages. Retries, limits and dead letters are your code.
  • PostgreSQL. SELECT ... FOR UPDATE SKIP LOCKED lets many workers take rows without blocking each other; the PostgreSQL docs say it suits queue-like tables. pgmq wraps this in SQS-style functions with a visibility timeout, a read count and an archive. Its README promises “exactly once” delivery within a visibility timeout; across timeouts it is at-least-once. For modest volumes next to data already in PostgreSQL, this is often the simplest answer.

Durability and replication

RabbitMQ quorum queues

NATS JetStream (R3)

Kafka

Replication

Raft per queue (Ra library), default 3 members

Raft group per stream, plus a meta group

Leader and follower replicas

Write is safe when

Publisher confirm received

PubAck received after quorum commit

Producer ack with acks=all and min.insync.replicas

Cluster metadata

Khepri (Raft), majority required since 4.3

Meta Raft group

KRaft controller quorum

The NATS clustering docs are unusually frank: a PubAck confirms quorum, not disk, and because JetStream syncs file storage every 2 minutes by default, two peers crashing within the same window can lose an acknowledged write. sync_interval: always closes the gap at a throughput cost, and 2.15.0 made it cheaper for replicated streams by syncing only the write-ahead log.

Operations: the part demos skip

Upgrades. RabbitMQ 4.3 upgrades only from 4.2.x, needs all stable feature flags enabled first and Erlang 27.0 or later, and mixed-version clusters should last hours, not days. The upgrade guide lets 3.13.x go straight to 4.2.x, then 4.3.x, so the Hyderabad team can move mirrored queues to quorum queues, then make two hops, or do a blue-green move with the tooling 4.2 added for 3.13 clusters. NATS rolling upgrades use lame-duck mode (SIGUSR2): a 10-second grace period, then client disconnects spread over 2 minutes by default. Kafka gates features through kafka-features.sh, so a binary upgrade alone may not enable share groups.

Alarms and flow control. RabbitMQ blocks publishers when a node crosses its memory high watermark (0.6 of RAM by default) or its free disk limit (50 MB by default; the production checklist calls that a development value and shows 4G as an example). Before that, flow control slows publishers that outrun their queues, and such connections show the flow state. JetStream’s limits default to 75% of RAM for memory storage and 75% of available disk under store_dir for file storage; pin max_file_store to what the volume really holds. NATS 2.11 added max_buffered_size and max_buffered_msgs to rate-limit ingest with a 429 reply.

Monitoring. RabbitMQ 4.2 renamed many Raft metrics, so dashboards on rabbitmq_raft* need updating. Watch dead letter depth and redelivery rates, not just queue length. In NATS, watch advisories and num_ack_pending per consumer; in Kafka, the share-partition lag added in 4.2.

Security advisories

All IDs below were confirmed as PUBLISHED through the CVE Services API on 8 October 2026.

Project

CVE

What happened

Fixed in

RabbitMQ

CVE-2026-57216

Loopback-only users such as guest could connect remotely through a PROXY-protocol path

3.13.15, 4.0.20, 4.1.11, 4.2.6

RabbitMQ

CVE-2026-57221

No authorisation check on passive queue.declare and exchange.declare

3.13.15, 4.0.20, 4.1.11, 4.2.6

RabbitMQ

CVE-2026-67420

OAuth token refresh kept revoked tags such as impersonator

4.2.10, 4.3.5

NATS

CVE-2025-30215

Missing access controls on some JetStream API requests across accounts

2.10.27, 2.11.1

NATS

CVE-2026-29785

Pre-auth panic on the leafnode port with compression

2.11.14, 2.12.5

NATS

CVE-2026-33217

Subject ACLs not applied in the $MQTT.> namespace

2.11.15, 2.12.6

Kafka

CVE-2026-35554

Producer buffer race could deliver messages to the wrong topic

3.9.2, 4.0.2, 4.1.2, 4.2.0

Kafka

CVE-2026-33557

Default JWT validator in 4.1.0 and 4.1.1 accepted unsigned tokens

4.1.2, 4.2.0

ActiveMQ Classic

CVE-2026-34197

Code injection through the Jolokia bridge in the web console

5.19.4, 6.2.3

Artemis

CVE-2026-27446

Unauthenticated Core protocol client could force federation to a rogue broker

Apache Artemis 2.52.0

LavinMQ

CVE-2026-25767

Policymaker users could create shovels across vhosts

2.6.8

Scale matters too. An NVD keyword search on 8 October 2026 returned 69 RabbitMQ server CVE records published since January 2025, 54 of them in September 2026, and 14 NATS server records published in 2026, most on 25 March. The OAuth fix above lands in 4.2.10, which I could not find as a public GitHub release, so for most community users the fix is 4.3.5 or later. The patterns: keep management UIs and leafnode, Core and Jolokia ports off untrusted networks, treat OAuth and PROXY-protocol setups as security-sensitive, and budget a patch upgrade every month.

Benchmarks: whose run is it?

Most numbers you will see are vendor runs. Confluent’s “Benchmarking RabbitMQ vs Kafka vs Pulsar” (21 August 2020) reports peak throughput of 605 MB/s for Kafka, 305 MB/s for Pulsar and 38 MB/s for mirrored RabbitMQ: Confluent’s own run, made before RabbitMQ 4.x. The RabbitMQ team’s AMQP 1.0 benchmark post (21 August 2024) reports 4.5 times lower throughput on 3.13 than on 4.0 for the same test: the RabbitMQ team’s own run. Neither tells you how your claims workload behaves with confirms, 4-minute jobs and retries.

A fair in-house test fixes the guarantees first (confirms, explicit acks, three replicas), uses real message sizes and processing times, kills consumers and nodes, and measures p99 latency and redelivery counts rather than peak throughput. nats bench, RabbitMQ PerfTest and kafka-share-consumer-perf-test.sh help.

A decision guide

Your situation

Reasonable first choice

Job queues with routing, retries, dead letters and priorities; existing AMQP clients

RabbitMQ 4.3 with quorum queues

Many small services, edge or multi-site, request-reply plus some persistence

NATS with JetStream

Kafka is already your platform and jobs are independent

Kafka share groups, with your own dead letter step

Large Java JMS estate

Apache Artemis or ActiveMQ Classic, patched

A few hundred jobs a second, data already in PostgreSQL

SKIP LOCKED or pgmq

Replay plus queue on one cluster, team willing to run BookKeeper

Pulsar, after reading the redelivery caveats

For the Hyderabad team, I would not add a broker yet. The claims pipeline belongs on RabbitMQ: move mirrored queues to quorum queues, upgrade through 4.2 to 4.3.x, set delivery limits, delayed retry and at-least-once dead lettering by policy, and audit consumers for nack used as “give up”. Analytics stays on Kafka. For the kiosks, run a pilot: NATS leafnodes suit hospital sites, but a third system to patch is a cost to write down first.

A practical checklist

  1. Write down the contract: retry limit, back-off, dead letter target, ordering need, maximum processing time.
  2. Make every consumer idempotent; all of these brokers redeliver.
  3. Set limits explicitly. Do not inherit 20, unlimited or 5 by accident.
  4. Use publisher confirms or PubAcks; an unconfirmed publish has failed.
  5. Give every queue a dead letter path and alert on its depth.
  6. Choose failure outcomes on purpose: reject or nack, term or nak, REJECT or RELEASE.
  7. Run three nodes, never two.
  8. Raise RabbitMQ’s free disk limit and pin JetStream storage limits.
  9. Rehearse the upgrade path on a copy.
  10. Put the broker on the monthly patch calendar.

Common mistakes

  • Using nack for poison messages on RabbitMQ 4.3. It no longer counts towards the delivery limit; use reject or a dead letter outcome.
  • Assuming JetStream drops failed messages. On a WorkQueue stream they stay after MaxDeliver, using storage, until you remove them.
  • Starting a Kafka share group and seeing nothing. The default start is the latest offset.
  • Setting AckWait or the lock duration shorter than the slowest job. Two workers then process the same claim.

Drill: RabbitMQ, NATS and Kafka on one VM

No root and no containers. On 8 October 2026, from about 2:25 PM to 2:45 PM IST, on a shared Linux VM with 8 vCPUs (Intel Xeon) and 15.6 GiB of RAM, I ran each broker on 127.0.0.1 only, one at a time, and stopped all processes afterwards.

  • RabbitMQ 4.3.6 generic Unix build (signature checked against the RabbitMQ release key) on Erlang/OTP 27.3.4.1 from Debian packages, unpacked into my own directory. Khepri was enabled, the memory watermark was 0.6 of RAM (about 10.07 GB) and the disk limit 50 MB. Client: pika 1.4.4.
  • NATS server 2.15.0 (SHA-256 matched SHA256SUMS). Client: nats-py 2.16.0.
  • Kafka 4.3.1 (SHA-512 matched the Apache file) as a single combined KRaft node on Temurin JDK 21.0.12.1.

RabbitMQ 4.3.6, quorum queues

Test

Result

basic.reject with requeue, x-delivery-limit 3

4 deliveries (x-delivery-count 1, 2, 3 on redeliveries), then dead-lettered with reason delivery_limit

basic.nack with requeue, x-delivery-limit 3

30 deliveries, never dead-lettered; still in the queue with x-acquired-count 30 and no x-delivery-count

x-consumer-timeout 5,000 ms, consumer never acks

Consumer cancelled at 5.0 s, message available again at 5.0 s with x-acquired-count 1

Delayed retry, min 1 s, max 4 s, limit 6, reject each time

Gaps of 1.01, 2.00, 3.01, 4.02, 4.02 and 4.02 s; dead-lettered after the 7th delivery

The policy shown above (limit 5, max 30 s)

6 deliveries, gaps of 1.02, 2.00, 3.01, 4.00 and 5.01 s, then in dlq.orders

NATS 2.15.0, JetStream

Test

Result

AckWait 2 s, MaxDeliver 3, worker never answers

Deliveries at 0.0, 2.0 and 4.0 s; one max-deliveries advisory at 6.01 s; message still in the WorkQueue stream

BackOff 1, 2, 4 s, MaxDeliver 4

Gaps of 1.0, 2.0 and 4.0 s

Plain nak, then nak with 1.5 s delay

Redelivered immediately, then after 1.501 s; term removed the message from the stream

Nats-TTL: 2s on a stream with message TTLs allowed

Message gone after 2.03 s

Snippet above

Message moved to DLQ only while a pull request was pending

Kafka 4.3.1, share groups

Test

Result

kafka-features.sh describe on a fresh cluster

share.version finalised at 1; maximum supported 1

--release on one record

Deliveries 1 to 5, then archived

New share group, default settings, 2,000 records already in the topic

0 records consumed in 20 s

Two share consumers on a 1-partition topic, 400 records produced after both joined

197 and 203 records, all 400 seen once, no duplicates; both assigned partition 0

In a first try with 2,000 records already waiting, the consumer that joined first took all of them: records are shared as consumers fetch, not split in advance.

Honest limits: every broker ran as a single node, so I tested no replication, failover, partitions or upgrades. Timings come from one client on a shared VM and show the broker’s retry schedule, not throughput or latency under load. Erlang came from Debian packages rather than RabbitMQ’s own builds, though 27.x is listed as supported for 4.3. I did not test Pulsar, Artemis, LavinMQ or pgmq.

What to unlearn and re-learn

  • Unlearn “RabbitMQ mirrored queues give us HA”. Re-learn quorum queues, publisher confirms and the majority rule that now also covers metadata.
  • Unlearn “nack means retry until the limit”. Re-learn that in RabbitMQ 4.3 only failures count, and choose reject, nack or release on purpose.
  • Unlearn “every broker has a dead letter queue”. Re-learn that JetStream and Kafka share groups leave that step to you.
  • Unlearn “the licence and the foundation never change”. Re-learn to read the LICENSE file and governance news on each upgrade; NATS came close to a licence change in 2025.

Revisit your message queue before you add another broker

The Hyderabad team does not need a fourth broker or a single winner. Learn what each broker does when a consumer fails: the counter, the timer, and where a poison message ends up. Unlearn defaults carried over from older versions. Re-learn RabbitMQ 4.3’s counting rules, JetStream’s advisories and Kafka’s share groups from the release notes, not from blog posts. Practise on one VM, as I did, with a message that always fails and a worker that never acks. Then apply the broker you already run with explicit limits, a dead letter path and a patch calendar, and revisit the choice when your workload, not a trend, asks for something else.

Sources

comments powered by Disqus

Releted Posts

Kafka and Redpanda: revisit the choice before you treat them as the same bus

Many teams say “we run Kafka” when they mean “our services talk to a Kafka-compatible broker”. That shortcut was useful when one open source broker dominated the protocol.

Read more

Vector search in 2026: revisit pgvector, Qdrant, Milvus, Weaviate before you add a vector database

An e-commerce company in Pune runs customer support on PostgreSQL. The database holds about 2 million help articles, product Q&A threads and resolved tickets, in English and a fair amount of Hinglish.

Read more

Service mesh in 2026: revisit Istio ambient, Linkerd and Cilium before you add sidecars

A logistics company in Bengaluru runs a production Kubernetes cluster with 14 nodes and about 420 pods. Three requests landed in one sprint.

Read more