Fluent Bit, Vector and the OpenTelemetry Collector: revisit your log pipeline before you ship more logs

A logistics company in Pune runs about 600 pods across three Kubernetes clusters. A Fluentd DaemonSet, set up in 2019 with a dozen Ruby plugins, ships about 25 crore log lines a day to Loki. During a festive sale last month, Loki was unhealthy for 40 minutes, and the next morning the team found a 25-minute hole in the order service logs exactly when they needed them. Now the SRE lead wants Fluent Bit, the platform team wants the OpenTelemetry Collector because the apps are moving to OpenTelemetry SDKs for traces, and one developer likes Vector because its transform language is easy to test.

All three are reasonable, and none is safe by default in the way the team assumes. The agent layer decides what happens to a log line when the backend is slow, down or rejecting data, and when the agent itself restarts. This post compares Fluent Bit, Vector and the OpenTelemetry Collector on that question, with notes on Fluentd, Logstash, Grafana Alloy and AxoSyslog. Loki, OpenSearch, ClickHouse and VictoriaLogs appear only as destinations; this is not a storage comparison. Facts come from release notes, official docs, licence files and CVE records, checked on 8 October 2026; the drill is my own run.

The short version

  • Fluent Bit’s default retry_limit is 1. In my drill, a 30-second outage of the destination lost 2,200 of 6,000 lines with default settings. With retry_limit: no_limits and filesystem buffering, nothing was lost, even when I killed the agent.
  • Vector blocks and applies backpressure by default, so an outage alone lost nothing. A kill -9 during the outage lost 1,492 lines, because the file checkpoint had moved past lines still held in memory. End-to-end acknowledgements fixed that with no duplicates.
  • The OpenTelemetry Collector retries for 5 minutes with a 1,000-request in-memory queue by default. A crash loses that queue, and file_log (formerly filelog) starts at the end of a file unless you give it storage. With storage on both sides, I still lost 20 lines in each of three runs.
  • OpenTelemetry became a CNCF graduated project on 11 May 2026. Collector releases in 2026 renamed core components (otlphttp is now otlp_http, filelog is now file_log); old names still work but log deprecation warnings.
  • Vector is maintained by Datadog and is still pre-1.0, with a minor release about every six weeks. Fluent Bit’s main corporate backer, Chronosphere, is now part of Palo Alto Networks; the project stays Apache 2.0 under CNCF.
  • Fluent Bit’s classic .conf format is due to be deprecated at the end of 2026. If your DaemonSet still uses it, plan the YAML move with your next upgrade.

Where things stand on 8 October 2026

Dates are from GitHub release pages, converted to IST.

Project

Latest release

What it is

Licence

Fluent Bit

5.1.3 on 1 October 2026; 5.0.10 on 4 September

C agent and aggregator for logs, metrics, traces and profiles

Apache 2.0

Vector

0.59.0 on 6 October 2026

Rust pipeline for logs and metrics, with VRL transforms

MPL 2.0

OpenTelemetry Collector

0.162.0 binaries on 29 September 2026 (core modules v1.68.0)

Go collector for traces, metrics, logs and profiles

Apache 2.0

Fluentd

1.19.4 on 28 September 2026

Ruby and C log collector with a large plugin ecosystem

Apache 2.0

Logstash

9.5.5, 9.4.8 and 8.19.23 on 6 October 2026

JRuby pipeline from Elastic

Apache 2.0 outside x-pack, Elastic License inside

Grafana Alloy

1.20.1 on 28 September 2026

OpenTelemetry Collector distribution with Prometheus pipelines

Apache 2.0

AxoSyslog

4.29.0 on 7 October 2026

syslog-ng based collector from Axoflow

GPL-3.0-or-later

Date

Event

9 June 2011

First Fluentd commit, per its CNCF project page

2014

Fluent Bit created at Treasure Data

8 November 2016

Fluentd accepted into CNCF at incubating level

11 April 2019

Fluentd graduates; Fluent Bit is graduated under the Fluentd umbrella

7 May 2019

OpenTelemetry, the merger of OpenTracing and OpenCensus, accepted into CNCF

11 February 2021

Datadog acquires Timber Technologies, the company behind Vector

28 November 2023

Collector release v1.0.0/v0.90.0 starts dual versioning: stable modules at 1.x, binaries still at 0.x

22 January 2024

Chronosphere acquires Calyptia, the company founded by the Fluent creators

9 April 2024

Grafana Alloy announced; Grafana Agent deprecated

1 November 2025

Grafana Agent reaches end of life

29 January 2026

Palo Alto Networks completes its acquisition of Chronosphere

24 March 2026

Fluent Bit 5.0.0

11 May 2026

OpenTelemetry moves to CNCF graduated level (announced 21 May)

6 August 2026

Fluent Bit 5.1.0 with multi-worker network inputs and an input rate gate

6 October 2026

Vector 0.59.0 drops Kubernetes 1.31 support

End of 2026

Fluent Bit classic .conf format due to be deprecated

Why the agent layer deserves a revisit

Most design effort goes into the backend; the earlier posts on Prometheus long-term storage and Elasticsearch and OpenSearch covered that side. The agent gets installed once from a Helm chart and forgotten.

Yet the agent decides three things on a bad day: where data waits when the destination cannot take it (memory, disk or nowhere), how long it keeps trying and what giving up means (drop, block the source, or move data aside), and what survives a restart of the agent itself, which in Kubernetes happens on every node drain and chart upgrade. The Pune team’s 25-minute hole was probably not a Loki problem but an agent that retried for a while, then discarded what it held.

How the three are built

Fluent Bit 5.1

Vector 0.59

OpenTelemetry Collector 0.162

Language

C

Rust

Go

Pipeline model

Inputs, parsers, filters, processors, outputs; routing by tag

Sources, transforms, sinks wired by inputs

Receivers, processors, exporters, connectors, per-signal pipelines

Transform language

Lua, processors, SQL stream processing

VRL (Vector Remap Language)

OTTL (OpenTelemetry Transformation Language)

Config format

YAML (standard since 3.2); classic .conf until end of 2026

YAML, TOML or JSON

YAML

Signals

Logs, metrics, traces, profiles

Logs and metrics (traces in a few components)

Traces, metrics, logs; profiles at alpha

Governance

CNCF graduated (under Fluentd); the creators’ company Calyptia is now part of Chronosphere (Palo Alto Networks)

Maintained by Datadog’s Community Open Source Engineering team

CNCF graduated; Collector SIG with many vendors

Fluent Bit

From embedded agent to the default in many clusters

The Fluent Bit docs say the Fluentd team at Treasure Data started it in 2014 for embedded Linux and gateways. Containers made it popular on ordinary servers. The project’s own comparison table puts Fluentd at more than 60 MB of memory and Fluent Bit at about 450 KB: the project’s own figure, not a sizing guide.

On ownership: Eduardo Silva and Anurag Gupta, the creators of the Fluent projects, founded Calyptia, which Chronosphere acquired on 22 January 2024. Palo Alto Networks agreed to buy Chronosphere on 19 November 2025 and completed the deal on 29 January 2026. The licence (Apache 2.0) and CNCF status are unchanged, but it helps to know who funds the people writing the roadmap.

What 5.0 and 5.1 changed

5.0.0 (March 2026) added worker threads to the shared HTTP server behind the http, splunk, elasticsearch, opentelemetry and prometheus_remote_write inputs, broader OTLP coverage, OAuth2 and backpressure duration metrics. 5.1.0 (August 2026) extended multi-worker listeners to tcp, udp, forward and syslog and added an input rate gate, adaptive flush, a FIPS mode and certificate reloading. 5.1.3 bounded concurrent flushes and retries under heavy backpressure and added tls.crl_file.

The throughput gains in those notes (OTLP ingestion “up to 4x faster” than 4.2; more than 2x with four listener workers in 5.1) are the Fluent Bit team’s own runs. The default is still one worker, so you only get them if you set workers or http_server.workers.

Buffering, retries and the default that bites

Records are buffered in chunks of about 2 MB, with three modes per input:

  • memory (the default): fast, lost on crash. When mem_buf_limit is reached the input is paused.
  • filesystem: chunks are written through mmap to storage.path and kept in memory when there is room. Set db on the tail input too, so read offsets survive restarts.
  • memrb, a memory ring buffer: when full, the oldest chunks are dropped instead of pausing the input, with memrb_dropped_chunks and memrb_dropped_bytes metrics.

Failed flushes are retried with exponential back-off and jitter (scheduler.base 5 s, scheduler.cap 2,000 s). The key line in the docs: Retry_Limit defaults to 1, and when a chunk exhausts its retries the data is discarded. So with defaults, an outage longer than one back-off interval loses data, as my drill showed. This is a likely cause of holes like the Pune team’s.

Two more traps: when an output’s storage.total_limit_size is reached, the oldest chunk is discarded; and a paused HTTP-based input accepts and immediately closes new connections, so senders must retry or lose data.

Fluent Bit also has a dead letter queue: with filesystem storage and storage.keep.rejected: on, chunks that hit the retry limit or a permanent error are copied to a rejected directory instead of vanishing. This config passes --dry-run on 5.1.3. In my drill, the same buffering and no_limits lost nothing, and in a separate run with retry_limit: 2 and the destination down, the chunk ended up in storage/rejected/:

service:
  flush: 1
  storage.path: /var/lib/fluent-bit/storage
  storage.sync: normal
  storage.keep.rejected: on
  storage.rejected.limit: 500M
pipeline:
  inputs:
    - name: tail
      path: /var/log/app/*.log
      db: /var/lib/fluent-bit/tail.db
      storage.type: filesystem
  outputs:
    - name: http
      match: '*'
      host: logs-gateway.internal
      port: 8080
      format: json
      retry_limit: no_limits
      storage.total_limit_size: 2G

Use no_limits on the main output and keep the dead letter queue for permanent errors such as a 400 from the backend.

Vector

Timber, Datadog and a long 0.x

Vector was built by Timber Technologies; Datadog announced the acquisition on 11 February 2021. The README says Vector is maintained by Datadog’s Community Open Source Engineering team, and the code is MPL 2.0. Development is active: 0.57.0 (15 July), 0.58.0 (26 August) and 0.59.0 (6 October 2026), roughly every six weeks. It is still pre-1.0, and the release notes advise stepping through minor versions because each can break things; 0.59.0 had six breaking changes, including removal of Kubernetes 1.31 support. Being company-run is not a problem in itself, but no foundation stands between the code and one vendor’s priorities.

Backpressure first, then buffers

Sinks get a default in-memory buffer of 500 events. When it is full, the default when_full: block pushes backpressure up to the source: a file source stops reading, a Kafka source stops consuming, an HTTP source slows its clients. drop_newest exists for push sources where loss is better than blocking clients.

A disk buffer is a write-ahead log on the sink, synced every 500 ms, with a minimum size of about 256 MiB. Vector stops itself on a disk I/O error such as a full volume rather than guess what reached disk, so monitor free space under data_dir.

The stronger guarantee is end-to-end acknowledgements. With acknowledgements.enabled: true on a sink, a participating source (file, Kafka, SQS, HTTP and others) commits a checkpoint or replies 200 only after the sink succeeds; Vector warns at start-up if a source, such as socket, cannot take part. In my drill this was the only Vector setting with no loss and no duplicates after a kill -9.

VRL

VRL is a small expression language for the remap transform. Errors from fallible functions must be handled, and Vector checks this at load time, so many mistakes show up before deployment. I tested this snippet on 0.59.0; it masks card-like numbers, parses key=value pairs and adds a field:

transforms:
  clean:
    type: remap
    inputs: [app_logs]
    source: |
      .message = redact(string!(.message), filters: [r'\b\d{12,19}\b'])
      . |= parse_key_value(.message) ?? {}
      .deployment_environment = "prod"
      del(.source_type)

For the input level=error seq=000042 card=4111111111111111 msg=payment-failed, the output had card set to [REDACTED] in both the message and the parsed field.

The OpenTelemetry Collector

Graduated, versioned in two lines

OpenTelemetry came from the merger of OpenTracing and OpenCensus, joined CNCF in May 2019 and graduated on 11 May 2026. The Collector uses two version lines: stable Go modules are 1.x (v1.68.0), while binaries and most components are 0.x (0.162.0). A distribution can reach 1.0 only once the core framework is stable, and each component has its own stability level per signal: development, alpha, beta, stable, deprecated or unmaintained.

The log path is uneven. From the metadata.yaml files on the main branch on 8 October 2026:

Component

Logs stability

In distributions

otlp receiver, otlp_grpc and otlp_http exporters

Stable

core, contrib, k8s, otlp

k8sattributes processor

Stable

contrib, k8s

batch processor, memory_limiter processor

Beta

core, contrib, k8s

file_log receiver, transform processor, file_storage extension

Beta

contrib, k8s

filter processor

Alpha

core, contrib, k8s

clickhouse exporter

Beta for logs

contrib

opensearch exporter

Alpha for logs

contrib

Core, contrib, k8s or your own

The project publishes otelcol (core), otelcol-contrib, otelcol-k8s, otelcol-otlp and an eBPF profiler build. Contrib holds every component at alpha or higher, and its README recommends limiting production Collectors to the components you need, for a smaller binary and attack surface, using the Collector Builder (ocb). The k8s distribution is a curated subset for Kubernetes. Vendors such as AWS, Elastic, Grafana, Splunk, Datadog and Bindplane ship their own distributions, listed but not endorsed by the project.

Renames you will see in the logs

2026 releases renamed components to snake_case: otlp and otlphttp became otlp_grpc and otlp_http in 0.144.0 (January), filelog became file_log in contrib 0.149.0 (March), and 0.158.0 added a queuebatch processor (renamed queue_batch in 0.162.0) to replace the legacy batch processor. Old names are deprecated aliases; my 0.162.0 runs started with old configs but logged a warning for each. Rename on your schedule, before an alias is removed.

Queues, retries and what a restart loses

Exporters use a shared helper. The defaults from its README: retry_on_failure enabled, starting at 5 s, capped at 30 s, giving up after 300 s (max_elapsed_time); sending_queue enabled with 1,000 requests and 10 consumers; block_on_overflow false, so data is rejected when the queue is full. Batching inside the queue is off unless you add batch: {}. Setting sending_queue.storage to a file_storage extension makes the queue persistent.

The receiver side matters as much. file_log defaults to start_at: end and keeps offsets in memory unless you set storage, so a restart without storage skips whatever was written while the Collector was down. Its README also warns that even with offset storage, logs can be dropped while moving through other components, which matches the 20 lines my drill lost. For a critical hop, the resiliency docs suggest a message queue such as Kafka between agent and gateway; the Kafka and Redpanda post covers that layer.

OTTL

OTTL statements are a function plus an optional where condition, run by the transform processor and others. This is what I tested on 0.162.0 with a key_value_parser operator on the receiver:

processors:
  transform:
    log_statements:
      - replace_pattern(log.body, "\\b\\d{12,19}\\b", "[REDACTED]")
      - set(log.attributes["card"], "[REDACTED]") where log.attributes["card"] != nil
      - set(log.severity_text, "ERROR") where log.attributes["level"] == "error"
      - set(resource.attributes["deployment.environment.name"], "prod")

The debug exporter showed the masked body, SeverityText: ERROR and the resource attribute.

The others, briefly

  • Fluentd is still maintained (release announcements come from ClearCode), and 1.19.4 goes into the fluent-package LTS 6.0.5. Its buffers retry for 72 hours by default (retry_timeout). The cost is Ruby memory use and a plugin ecosystem in which every gem is code you must patch. CNCF published a Fluentd to Fluent Bit migration guide in October 2025.
  • Logstash makes sense where Elastic is already the platform. Persistent queues are off by default (queue.type: memory), and the docs say they cannot protect inputs without a request-response protocol, such as TCP and UDP.
  • Grafana Alloy replaced Grafana Agent, which reached end of life on 1 November 2025. Alloy is an OpenTelemetry Collector distribution with Prometheus pipelines and its own block-based syntax, a natural fit for Grafana stacks.
  • AxoSyslog, Axoflow’s syslog-ng based collector (GPL-3.0-or-later), suits syslog-heavy network and security estates.
  • Elastic Agent and Beats (9.5.5 on 6 October 2026) and vendor Collector distributions are fine when you are committed to that vendor; check how far they are from upstream.

Delivery behaviour, side by side

Question

Fluent Bit 5.1

Vector 0.59

OTel Collector 0.162

Default buffer

Memory per input, paused at mem_buf_limit (no limit by default)

Memory, 500 events per sink

In-memory queue, 1,000 requests per exporter

When the buffer is full

Input paused

Block (backpressure)

Reject new data

Default retry limit

1 retry, then discard

Effectively unlimited for the http sink

300 s, then drop

Disk option

Filesystem buffering per input

Disk buffer per sink, at least about 256 MiB

file_storage for the queue and for file_log offsets

Destination ack back to the source

No

End-to-end acknowledgements

Not by default; wait_for_result makes the caller wait (I did not test it)

What a crash loses by default

Memory chunks; tail offsets unless db is set

Buffered events after the file checkpoint moved

Queue contents, plus lines written while down (start_at: end)

Dead letter path

storage.keep.rejected

None built in

None built in

All three are at-least-once at best. Duplicates after a restart are normal: Vector with a disk buffer and no acknowledgements sent 89 duplicates after a kill -9 in my drill. If duplicates matter, deduplicate in the backend using a stable event ID.

Kubernetes: DaemonSet, gateway, or both

Pattern

What it does

Watch out for

DaemonSet agent

Tails /var/log/pods on each node, adds pod metadata, forwards

Buffer on a hostPath volume, or a pod restart loses it; agent CPU and memory limits on busy nodes

Gateway Deployment

Receives OTLP or forward traffic, does heavy processing, holds credentials, fans out to backends

Needs its own persistent volume and retries; one replica for cluster-level receivers such as k8s_cluster to avoid duplicates

Agent, then Kafka, then gateway

Durable hop between tiers

One more system to run and patch

The OpenTelemetry Kubernetes docs prefer a DaemonSet for file_log and kubeletstats, and warn that cluster-level receivers in a DaemonSet or multi-replica Deployment duplicate data. Tail sampling needs all spans of a trace on one Collector, which is what the loadbalancing exporter is for. A sensible layout: a small agent per node that only collects, enriches and forwards, and a few gateway replicas that redact, sample and route with persistent queues.

Give the agent a memory limit that matches its buffer settings: a Collector without memory_limiter, or a Vector sink with a large in-memory buffer, can be OOM-killed under backpressure instead of slowing down.

OTLP and semantic conventions

OTLP is now the common language: Fluent Bit has native OTLP input and output, Vector has an opentelemetry source and sink, and the Collector is built around it. The latest releases are OTLP protobuf 1.11.1 (29 September 2026), specification 1.61.0 and semantic conventions 1.44.0 (5 August 2026).

Speaking OTLP is not the same as agreeing on field names. Semantic conventions define attributes such as service.name, k8s.pod.name and deployment.environment.name (the older deployment.environment is deprecated). If one cluster’s agent writes kubernetes.pod_name and another’s writes k8s.pod.name, queries break across clusters. Pick one scheme, set it at the agent, and test it with a sample record.

Resource use and benchmarks: whose run is it?

Most numbers come from the projects themselves. Fluent Bit’s release notes report its own gains; Vector’s README says “up to 10x faster than every alternative”, the Vector team’s own claim. A CNCF case study (12 November 2025) describes OpenAI cutting Fluent Bit CPU use by 50%: one user’s environment, not a comparison.

In my drill, only as a sense of scale, the highest sampled resident memory at 100 lines a second was about 19 to 21 MB for Fluent Bit, 52 to 61 MB for Vector and 221 to 226 MB for otelcol-contrib; the binaries were 76 MB, 156 MB and 407 MB. A slim ocb build would be smaller, but I did not measure one. A fair test uses your real lines, parsing rules and backend, and measures CPU per lakh lines, memory during a 30-minute outage, and loss and duplicates after a node drain.

Security advisories

All IDs below were confirmed as PUBLISHED through the CVE Services API on 8 October 2026.

Project

CVE

What happened

Fixed in

Fluent Bit

CVE-2025-12969

in_forward did not enforce security.users authentication in some configurations

4.0.13, 4.1.1 (per the Fluent Bit advisory)

Fluent Bit

CVE-2025-12972

out_file built file names from untrusted tags, allowing path traversal

4.0.13, 4.1.1

Fluent Bit

CVE-2025-12977

HTTP, Splunk and Elasticsearch inputs accepted unsanitised tag_key values

4.0.13, 4.1.1

Fluent Bit

CVE-2026-61674

Stack overflow in out_forward Secure Forward handshake with a malicious server

5.0.8

Fluentd

CVE-2026-44024

${tag} path traversal in out_file and similar outputs

1.19.3 (1.19.4 completes the fix)

Fluentd

CVE-2026-44025

in_monitor_agent exposed internal variables that may contain credentials

1.19.3

Fluentd

CVE-2026-44160

Compressed payloads to in_http and in_forward could exhaust memory

1.19.3

Vector

CVE-2026-77619, CVE-2026-77620

logstash source on 0.0.0.0:5044 could be crashed by oversized or nested compressed frames

0.57.0

Vector

CVE-2026-77621

file sink path templates allowed traversal from event fields

0.57.0

OTel Collector contrib

CVE-2026-42602

azure_auth extension authentication bypass

0.151.0

OTel Collector contrib

CVE-2026-55701

github receiver ignored required headers on webhooks

0.151.0

Logstash

CVE-2026-33466

Archive path traversal through the GeoIP database downloader

8.19.14, 9.2.8, 9.3.3

Grafana Alloy

CVE-2026-75889

prometheus.operator.servicemonitors could read local files and send them as a bearer token

1.19.0

The patterns repeat: listeners that trust whoever connects, tags or event fields used to build paths or URLs, and decompression without limits. Bind inputs to the network you mean, require TLS and authentication on forward and OTLP ports, never build output paths from tags you did not set, and run the agent as non-root. The Fluentd 1.19.4 notes say the 1.19.3 fix for CVE-2026-44024 was incomplete: take the latest patch, not the first fixed one.

Upgrade notes

  • Fluent Bit 4.x to 5.1. In 5.0, fluentbit_hot_reloaded_times became a counter and HTTP input settings moved to http_server.* (old names still accepted). In 5.1, Debian and Ubuntu package upgrades restart a running service. Move from .conf to YAML before the end-of-2026 deprecation.
  • Vector. Read every minor version’s upgrade guide; a jump from 0.50 to 0.59 means nine of them.
  • OpenTelemetry Collector. Releases come about every two weeks. Read breaking changes for the components you use, rename deprecated aliases, and keep core and contrib versions in step in ocb builds.
  • Fluentd. Stay on the latest 1.19.x or fluent-package LTS, and audit gems: fluent-plugin-s3 (CVE-2026-44162) and fluent-plugin-opentelemetry (CVE-2026-44163) had advisories in September 2026.
  • Grafana Agent. Anything still running it is unpatched; move it to Alloy.

A decision guide

Your situation

Reasonable first choice

Kubernetes node agent, mostly logs, tight memory

Fluent Bit 5.1 with YAML, filesystem buffering and no_limits

Apps moving to OpenTelemetry SDKs; traces, metrics and logs together

OpenTelemetry Collector, slim build as agent plus gateway

Heavy parsing, redaction and routing with tested transforms

Vector as an aggregator with end-to-end acknowledgements

Grafana stack, many Prometheus scrape jobs

Grafana Alloy

Large Fluentd estate with custom Ruby plugins

Keep Fluentd patched and move node collection to Fluent Bit first

Elastic-centred platform

Elastic Agent or Logstash with persistent queues

Syslog from network and security devices

AxoSyslog or Fluent Bit’s syslog input, behind a gateway

For the Pune team, I would not rip everything out at once. Replace the Fluentd DaemonSet with Fluent Bit 5.1 in YAML, with filesystem buffering on a hostPath, retry_limit: no_limits and the dead letter queue, forwarding over OTLP to two or three Collector gateway replicas that redact with OTTL and keep a persistent queue sized for an hour of traffic. The apps’ SDKs send traces to the same gateways. Vector is a good gateway too if the team prefers VRL and acknowledgements; what matters more is writing down which hop owns durability, and testing it.

A practical checklist

  1. Write down the loss budget: how long an outage must be survived, and whether duplicates are acceptable.
  2. Check every retry default in the agent you run. Fluent Bit’s is 1; the Collector’s is 5 minutes.
  3. Put buffers on disk that survives a pod restart, and size it in hours of traffic.
  4. Persist read positions: Fluent Bit db, Collector file_log storage, Vector data_dir.
  5. Prefer acknowledgement or backpressure over silent drop wherever the source can wait.
  6. Alert on drops, not only on errors: retry exhaustion, queue full, dropped chunks, enqueue failures.
  7. Standardise attribute names with semantic conventions at the agent.
  8. Build a slim Collector with ocb or use the k8s distribution instead of contrib in production.
  9. Expose listeners deliberately, with TLS and authentication, and patch monthly.
  10. Rehearse an outage: stop the backend for 30 minutes and drain a node, then count what arrived.

Common mistakes

  • Trusting Fluent Bit’s defaults for an outage. One retry and the chunk is discarded.
  • Running the Collector’s file_log without storage. A restart skips everything written while it was down.
  • Using emptyDir for buffers. It disappears with the pod; mount the buffer from the node.
  • Shipping otelcol-contrib to every node. It is a 407 MB binary of components you mostly do not use, each one attack surface.

Drill: losing logs on purpose on one VM

On 8 October 2026, from about 6:30 PM to 6:40 PM IST, on a shared Linux VM with 8 vCPUs (Intel Xeon) and 15.6 GiB of RAM, without root or containers, I ran each agent on 127.0.0.1 only and stopped every process afterwards.

  • Fluent Bit 5.1.3 from the official Debian trixie package (SHA-256 matched the repository Packages index, whose InRelease file had a good signature from the Fluent Bit release key), unpacked into my own directory with libyaml from Debian.
  • Vector 0.59.0 x86_64-unknown-linux-musl (SHA-256 matched SHA256SUMS).
  • otelcol-contrib 0.162.0 for linux_amd64 (SHA-256 matched the published .sha256 file).

Method: a script wrote 6,000 numbered lines at 100 lines a second to a file; each agent tailed it and posted to a small Python HTTP receiver that recorded sequence numbers. The receiver was down from second 10 to second 40. In crash runs, the agent got SIGKILL at second 25 and was restarted at second 30, with the receiver still down.

Agent and settings

Crash?

Unique lines received of 6,000

Duplicates

Lost

Fluent Bit, defaults (memory, retry_limit 1)

No

3,800

0

2,200 (lines 973 to 3172); 22 “cannot be retried” log lines

Fluent Bit, memory, retry_limit: no_limits

No

6,000

0

0

Fluent Bit, filesystem buffering, db, no_limits

Yes

6,000

0

0

Vector, defaults

No

6,000

0

0

Vector, defaults

Yes

4,508

0

1,492 (lines 916 to 2407)

Vector, end-to-end acknowledgements, memory buffer

Yes

6,000

0

0

Vector, disk buffer (minimum size), no acknowledgements, two runs

Yes

6,000

89 each run

0

Collector, defaults

No

6,000

0

0

Collector, defaults

Yes

3,940

0

2,060 (lines 972 to 3031)

Collector, file_storage for file_log and sending_queue, three runs

Yes

5,980

0

20 each run (around lines 2475 to 2495)

Reading the results: Fluent Bit lost the chunks whose single retry fell inside the outage, and kept those whose retry came after the receiver returned. Vector’s loss is consistent with its file checkpoint having moved past lines that were still in memory. The Collector with defaults lost its in-memory queue at the kill plus the lines written while it was down, since file_log restarted at the end of the file. The 20 lines lost with storage sat just before the kill each time, consistent with the README’s warning about drops between components. In a separate test, Fluent Bit with retry_limit: 2 and storage.keep.rejected: on put the 50-line chunk it could not deliver into storage/rejected/ instead of dropping it.

Honest limits: one VM, one file, short lines, 100 lines a second, a 30-second outage and one kill per run. I did not test Kubernetes, log rotation, multiline parsing, TLS, real backends, node drains or high throughput. Memory figures are the highest sampled VmRSS during these runs, not a benchmark. Each setting was run once unless marked otherwise. I did not test Fluentd, Logstash, Alloy or AxoSyslog.

What to unlearn and re-learn

  • Unlearn “the agent retries, so logs are safe”. Re-learn each agent’s retry limit and what happens when it is reached: discard, drop, or a dead letter queue.
  • Unlearn “a disk buffer means no loss and no duplicates”. Re-learn that read positions, buffers and acknowledgements are three separate things, and you need all three to line up.
  • Unlearn “OpenTelemetry means one stable thing”. Re-learn that each Collector component has its own stability level per signal, and that names still change between releases.
  • Unlearn “this project will always be run the way it is today”. Re-learn who maintains each agent: a CNCF community, a single vendor, or a vendor that was just acquired.

Revisit your log pipeline before you ship more logs

The Pune team does not need the fastest agent; it needs one whose failure behaviour it understands. Learn where each agent keeps data when the backend says no, and for how long. Unlearn defaults inherited from a Helm chart written years ago. Re-learn buffering, read positions and acknowledgements from the current docs and release notes. Practise with a stopped backend and a killed agent, as I did on one VM, and count the lines. Then apply the result: a durable hop you can name, a loss budget you have tested, and a patch calendar for the agent on every node. Revisit the choice when your signals, your backends or the projects’ owners change.

Sources

comments powered by Disqus

Releted Posts

Argo CD and Flux: revisit your GitOps controller before you scale or migrate

A platform team in Pune looks after 14 Kubernetes clusters. Most of them are managed by one central Argo CD instance that the team set up in 2021.

Read more

Ingress NGINX retirement: revisit your Kubernetes ingress before you migrate

A platform team plans its Kubernetes 1.37 upgrade for the next sprint. The cluster checklist is green, except one line nobody owns: the ingress-nginx Helm chart, pinned at 4.

Read more

Service mesh in 2026: revisit Istio ambient, Linkerd and Cilium before you add sidecars

A logistics company in Bengaluru runs a production Kubernetes cluster with 14 nodes and about 420 pods. Three requests landed in one sprint.

Read more