Fluent Bit, Vector and the OpenTelemetry Collector: revisit your log pipeline before you ship more logs
A logistics company in Pune runs about 600 pods across three Kubernetes clusters. A Fluentd DaemonSet, set up in 2019 with a dozen Ruby plugins, ships about 25 crore log lines a day to Loki. During a festive sale last month, Loki was unhealthy for 40 minutes, and the next morning the team found a 25-minute hole in the order service logs exactly when they needed them. Now the SRE lead wants Fluent Bit, the platform team wants the OpenTelemetry Collector because the apps are moving to OpenTelemetry SDKs for traces, and one developer likes Vector because its transform language is easy to test.
All three are reasonable, and none is safe by default in the way the team assumes. The agent layer decides what happens to a log line when the backend is slow, down or rejecting data, and when the agent itself restarts. This post compares Fluent Bit, Vector and the OpenTelemetry Collector on that question, with notes on Fluentd, Logstash, Grafana Alloy and AxoSyslog. Loki, OpenSearch, ClickHouse and VictoriaLogs appear only as destinations; this is not a storage comparison. Facts come from release notes, official docs, licence files and CVE records, checked on 8 October 2026; the drill is my own run.
The short version
- Fluent Bit’s default
retry_limitis 1. In my drill, a 30-second outage of the destination lost 2,200 of 6,000 lines with default settings. Withretry_limit: no_limitsand filesystem buffering, nothing was lost, even when I killed the agent. - Vector blocks and applies backpressure by default, so an outage alone lost nothing. A
kill -9during the outage lost 1,492 lines, because the file checkpoint had moved past lines still held in memory. End-to-end acknowledgements fixed that with no duplicates. - The OpenTelemetry Collector retries for 5 minutes with a 1,000-request in-memory queue by default. A crash loses that queue, and
file_log(formerlyfilelog) starts at the end of a file unless you give it storage. With storage on both sides, I still lost 20 lines in each of three runs. - OpenTelemetry became a CNCF graduated project on 11 May 2026. Collector releases in 2026 renamed core components (
otlphttpis nowotlp_http,filelogis nowfile_log); old names still work but log deprecation warnings. - Vector is maintained by Datadog and is still pre-1.0, with a minor release about every six weeks. Fluent Bit’s main corporate backer, Chronosphere, is now part of Palo Alto Networks; the project stays Apache 2.0 under CNCF.
- Fluent Bit’s classic
.confformat is due to be deprecated at the end of 2026. If your DaemonSet still uses it, plan the YAML move with your next upgrade.
Where things stand on 8 October 2026
Dates are from GitHub release pages, converted to IST.
Project
Latest release
What it is
Licence
Fluent Bit
5.1.3 on 1 October 2026; 5.0.10 on 4 September
C agent and aggregator for logs, metrics, traces and profiles
Apache 2.0
Vector
0.59.0 on 6 October 2026
Rust pipeline for logs and metrics, with VRL transforms
MPL 2.0
OpenTelemetry Collector
0.162.0 binaries on 29 September 2026 (core modules v1.68.0)
Go collector for traces, metrics, logs and profiles
Apache 2.0
Fluentd
1.19.4 on 28 September 2026
Ruby and C log collector with a large plugin ecosystem
Apache 2.0
Logstash
9.5.5, 9.4.8 and 8.19.23 on 6 October 2026
JRuby pipeline from Elastic
Apache 2.0 outside x-pack, Elastic License inside
Grafana Alloy
1.20.1 on 28 September 2026
OpenTelemetry Collector distribution with Prometheus pipelines
Apache 2.0
AxoSyslog
4.29.0 on 7 October 2026
syslog-ng based collector from Axoflow
GPL-3.0-or-later
Date
Event
9 June 2011
First Fluentd commit, per its CNCF project page
2014
Fluent Bit created at Treasure Data
8 November 2016
Fluentd accepted into CNCF at incubating level
11 April 2019
Fluentd graduates; Fluent Bit is graduated under the Fluentd umbrella
7 May 2019
OpenTelemetry, the merger of OpenTracing and OpenCensus, accepted into CNCF
11 February 2021
Datadog acquires Timber Technologies, the company behind Vector
28 November 2023
Collector release v1.0.0/v0.90.0 starts dual versioning: stable modules at 1.x, binaries still at 0.x
22 January 2024
Chronosphere acquires Calyptia, the company founded by the Fluent creators
9 April 2024
Grafana Alloy announced; Grafana Agent deprecated
1 November 2025
Grafana Agent reaches end of life
29 January 2026
Palo Alto Networks completes its acquisition of Chronosphere
24 March 2026
Fluent Bit 5.0.0
11 May 2026
OpenTelemetry moves to CNCF graduated level (announced 21 May)
6 August 2026
Fluent Bit 5.1.0 with multi-worker network inputs and an input rate gate
6 October 2026
Vector 0.59.0 drops Kubernetes 1.31 support
End of 2026
Fluent Bit classic .conf format due to be deprecated
Why the agent layer deserves a revisit
Most design effort goes into the backend; the earlier posts on Prometheus long-term storage and Elasticsearch and OpenSearch covered that side. The agent gets installed once from a Helm chart and forgotten.
Yet the agent decides three things on a bad day: where data waits when the destination cannot take it (memory, disk or nowhere), how long it keeps trying and what giving up means (drop, block the source, or move data aside), and what survives a restart of the agent itself, which in Kubernetes happens on every node drain and chart upgrade. The Pune team’s 25-minute hole was probably not a Loki problem but an agent that retried for a while, then discarded what it held.
How the three are built
Fluent Bit 5.1
Vector 0.59
OpenTelemetry Collector 0.162
Language
C
Rust
Go
Pipeline model
Inputs, parsers, filters, processors, outputs; routing by tag
Sources, transforms, sinks wired by inputs
Receivers, processors, exporters, connectors, per-signal pipelines
Transform language
Lua, processors, SQL stream processing
VRL (Vector Remap Language)
OTTL (OpenTelemetry Transformation Language)
Config format
YAML (standard since 3.2); classic .conf until end of 2026
YAML, TOML or JSON
YAML
Signals
Logs, metrics, traces, profiles
Logs and metrics (traces in a few components)
Traces, metrics, logs; profiles at alpha
Governance
CNCF graduated (under Fluentd); the creators’ company Calyptia is now part of Chronosphere (Palo Alto Networks)
Maintained by Datadog’s Community Open Source Engineering team
CNCF graduated; Collector SIG with many vendors
Fluent Bit
From embedded agent to the default in many clusters
The Fluent Bit docs say the Fluentd team at Treasure Data started it in 2014 for embedded Linux and gateways. Containers made it popular on ordinary servers. The project’s own comparison table puts Fluentd at more than 60 MB of memory and Fluent Bit at about 450 KB: the project’s own figure, not a sizing guide.
On ownership: Eduardo Silva and Anurag Gupta, the creators of the Fluent projects, founded Calyptia, which Chronosphere acquired on 22 January 2024. Palo Alto Networks agreed to buy Chronosphere on 19 November 2025 and completed the deal on 29 January 2026. The licence (Apache 2.0) and CNCF status are unchanged, but it helps to know who funds the people writing the roadmap.
What 5.0 and 5.1 changed
5.0.0 (March 2026) added worker threads to the shared HTTP server behind the http, splunk, elasticsearch, opentelemetry and prometheus_remote_write inputs, broader OTLP coverage, OAuth2 and backpressure duration metrics. 5.1.0 (August 2026) extended multi-worker listeners to tcp, udp, forward and syslog and added an input rate gate, adaptive flush, a FIPS mode and certificate reloading. 5.1.3 bounded concurrent flushes and retries under heavy backpressure and added tls.crl_file.
The throughput gains in those notes (OTLP ingestion “up to 4x faster” than 4.2; more than 2x with four listener workers in 5.1) are the Fluent Bit team’s own runs. The default is still one worker, so you only get them if you set workers or http_server.workers.
Buffering, retries and the default that bites
Records are buffered in chunks of about 2 MB, with three modes per input:
memory(the default): fast, lost on crash. Whenmem_buf_limitis reached the input is paused.filesystem: chunks are written throughmmaptostorage.pathand kept in memory when there is room. Setdbon thetailinput too, so read offsets survive restarts.memrb, a memory ring buffer: when full, the oldest chunks are dropped instead of pausing the input, withmemrb_dropped_chunksandmemrb_dropped_bytesmetrics.
Failed flushes are retried with exponential back-off and jitter (scheduler.base 5 s, scheduler.cap 2,000 s). The key line in the docs: Retry_Limit defaults to 1, and when a chunk exhausts its retries the data is discarded. So with defaults, an outage longer than one back-off interval loses data, as my drill showed. This is a likely cause of holes like the Pune team’s.
Two more traps: when an output’s storage.total_limit_size is reached, the oldest chunk is discarded; and a paused HTTP-based input accepts and immediately closes new connections, so senders must retry or lose data.
Fluent Bit also has a dead letter queue: with filesystem storage and storage.keep.rejected: on, chunks that hit the retry limit or a permanent error are copied to a rejected directory instead of vanishing. This config passes --dry-run on 5.1.3. In my drill, the same buffering and no_limits lost nothing, and in a separate run with retry_limit: 2 and the destination down, the chunk ended up in storage/rejected/:
service:
flush: 1
storage.path: /var/lib/fluent-bit/storage
storage.sync: normal
storage.keep.rejected: on
storage.rejected.limit: 500M
pipeline:
inputs:
- name: tail
path: /var/log/app/*.log
db: /var/lib/fluent-bit/tail.db
storage.type: filesystem
outputs:
- name: http
match: '*'
host: logs-gateway.internal
port: 8080
format: json
retry_limit: no_limits
storage.total_limit_size: 2G
Use no_limits on the main output and keep the dead letter queue for permanent errors such as a 400 from the backend.
Vector
Timber, Datadog and a long 0.x
Vector was built by Timber Technologies; Datadog announced the acquisition on 11 February 2021. The README says Vector is maintained by Datadog’s Community Open Source Engineering team, and the code is MPL 2.0. Development is active: 0.57.0 (15 July), 0.58.0 (26 August) and 0.59.0 (6 October 2026), roughly every six weeks. It is still pre-1.0, and the release notes advise stepping through minor versions because each can break things; 0.59.0 had six breaking changes, including removal of Kubernetes 1.31 support. Being company-run is not a problem in itself, but no foundation stands between the code and one vendor’s priorities.
Backpressure first, then buffers
Sinks get a default in-memory buffer of 500 events. When it is full, the default when_full: block pushes backpressure up to the source: a file source stops reading, a Kafka source stops consuming, an HTTP source slows its clients. drop_newest exists for push sources where loss is better than blocking clients.
A disk buffer is a write-ahead log on the sink, synced every 500 ms, with a minimum size of about 256 MiB. Vector stops itself on a disk I/O error such as a full volume rather than guess what reached disk, so monitor free space under data_dir.
The stronger guarantee is end-to-end acknowledgements. With acknowledgements.enabled: true on a sink, a participating source (file, Kafka, SQS, HTTP and others) commits a checkpoint or replies 200 only after the sink succeeds; Vector warns at start-up if a source, such as socket, cannot take part. In my drill this was the only Vector setting with no loss and no duplicates after a kill -9.
VRL
VRL is a small expression language for the remap transform. Errors from fallible functions must be handled, and Vector checks this at load time, so many mistakes show up before deployment. I tested this snippet on 0.59.0; it masks card-like numbers, parses key=value pairs and adds a field:
transforms:
clean:
type: remap
inputs: [app_logs]
source: |
.message = redact(string!(.message), filters: [r'\b\d{12,19}\b'])
. |= parse_key_value(.message) ?? {}
.deployment_environment = "prod"
del(.source_type)
For the input level=error seq=000042 card=4111111111111111 msg=payment-failed, the output had card set to [REDACTED] in both the message and the parsed field.
The OpenTelemetry Collector
Graduated, versioned in two lines
OpenTelemetry came from the merger of OpenTracing and OpenCensus, joined CNCF in May 2019 and graduated on 11 May 2026. The Collector uses two version lines: stable Go modules are 1.x (v1.68.0), while binaries and most components are 0.x (0.162.0). A distribution can reach 1.0 only once the core framework is stable, and each component has its own stability level per signal: development, alpha, beta, stable, deprecated or unmaintained.
The log path is uneven. From the metadata.yaml files on the main branch on 8 October 2026:
Component
Logs stability
In distributions
otlp receiver, otlp_grpc and otlp_http exporters
Stable
core, contrib, k8s, otlp
k8sattributes processor
Stable
contrib, k8s
batch processor, memory_limiter processor
Beta
core, contrib, k8s
file_log receiver, transform processor, file_storage extension
Beta
contrib, k8s
filter processor
Alpha
core, contrib, k8s
clickhouse exporter
Beta for logs
contrib
opensearch exporter
Alpha for logs
contrib
Core, contrib, k8s or your own
The project publishes otelcol (core), otelcol-contrib, otelcol-k8s, otelcol-otlp and an eBPF profiler build. Contrib holds every component at alpha or higher, and its README recommends limiting production Collectors to the components you need, for a smaller binary and attack surface, using the Collector Builder (ocb). The k8s distribution is a curated subset for Kubernetes. Vendors such as AWS, Elastic, Grafana, Splunk, Datadog and Bindplane ship their own distributions, listed but not endorsed by the project.
Renames you will see in the logs
2026 releases renamed components to snake_case: otlp and otlphttp became otlp_grpc and otlp_http in 0.144.0 (January), filelog became file_log in contrib 0.149.0 (March), and 0.158.0 added a queuebatch processor (renamed queue_batch in 0.162.0) to replace the legacy batch processor. Old names are deprecated aliases; my 0.162.0 runs started with old configs but logged a warning for each. Rename on your schedule, before an alias is removed.
Queues, retries and what a restart loses
Exporters use a shared helper. The defaults from its README: retry_on_failure enabled, starting at 5 s, capped at 30 s, giving up after 300 s (max_elapsed_time); sending_queue enabled with 1,000 requests and 10 consumers; block_on_overflow false, so data is rejected when the queue is full. Batching inside the queue is off unless you add batch: {}. Setting sending_queue.storage to a file_storage extension makes the queue persistent.
The receiver side matters as much. file_log defaults to start_at: end and keeps offsets in memory unless you set storage, so a restart without storage skips whatever was written while the Collector was down. Its README also warns that even with offset storage, logs can be dropped while moving through other components, which matches the 20 lines my drill lost. For a critical hop, the resiliency docs suggest a message queue such as Kafka between agent and gateway; the Kafka and Redpanda post covers that layer.
OTTL
OTTL statements are a function plus an optional where condition, run by the transform processor and others. This is what I tested on 0.162.0 with a key_value_parser operator on the receiver:
processors:
transform:
log_statements:
- replace_pattern(log.body, "\\b\\d{12,19}\\b", "[REDACTED]")
- set(log.attributes["card"], "[REDACTED]") where log.attributes["card"] != nil
- set(log.severity_text, "ERROR") where log.attributes["level"] == "error"
- set(resource.attributes["deployment.environment.name"], "prod")
The debug exporter showed the masked body, SeverityText: ERROR and the resource attribute.
The others, briefly
- Fluentd is still maintained (release announcements come from ClearCode), and 1.19.4 goes into the
fluent-packageLTS 6.0.5. Its buffers retry for 72 hours by default (retry_timeout). The cost is Ruby memory use and a plugin ecosystem in which every gem is code you must patch. CNCF published a Fluentd to Fluent Bit migration guide in October 2025. - Logstash makes sense where Elastic is already the platform. Persistent queues are off by default (
queue.type: memory), and the docs say they cannot protect inputs without a request-response protocol, such as TCP and UDP. - Grafana Alloy replaced Grafana Agent, which reached end of life on 1 November 2025. Alloy is an OpenTelemetry Collector distribution with Prometheus pipelines and its own block-based syntax, a natural fit for Grafana stacks.
- AxoSyslog, Axoflow’s syslog-ng based collector (GPL-3.0-or-later), suits syslog-heavy network and security estates.
- Elastic Agent and Beats (9.5.5 on 6 October 2026) and vendor Collector distributions are fine when you are committed to that vendor; check how far they are from upstream.
Delivery behaviour, side by side
Question
Fluent Bit 5.1
Vector 0.59
OTel Collector 0.162
Default buffer
Memory per input, paused at mem_buf_limit (no limit by default)
Memory, 500 events per sink
In-memory queue, 1,000 requests per exporter
When the buffer is full
Input paused
Block (backpressure)
Reject new data
Default retry limit
1 retry, then discard
Effectively unlimited for the http sink
300 s, then drop
Disk option
Filesystem buffering per input
Disk buffer per sink, at least about 256 MiB
file_storage for the queue and for file_log offsets
Destination ack back to the source
No
End-to-end acknowledgements
Not by default; wait_for_result makes the caller wait (I did not test it)
What a crash loses by default
Memory chunks; tail offsets unless db is set
Buffered events after the file checkpoint moved
Queue contents, plus lines written while down (start_at: end)
Dead letter path
storage.keep.rejected
None built in
None built in
All three are at-least-once at best. Duplicates after a restart are normal: Vector with a disk buffer and no acknowledgements sent 89 duplicates after a kill -9 in my drill. If duplicates matter, deduplicate in the backend using a stable event ID.
Kubernetes: DaemonSet, gateway, or both
Pattern
What it does
Watch out for
DaemonSet agent
Tails /var/log/pods on each node, adds pod metadata, forwards
Buffer on a hostPath volume, or a pod restart loses it; agent CPU and memory limits on busy nodes
Gateway Deployment
Receives OTLP or forward traffic, does heavy processing, holds credentials, fans out to backends
Needs its own persistent volume and retries; one replica for cluster-level receivers such as k8s_cluster to avoid duplicates
Agent, then Kafka, then gateway
Durable hop between tiers
One more system to run and patch
The OpenTelemetry Kubernetes docs prefer a DaemonSet for file_log and kubeletstats, and warn that cluster-level receivers in a DaemonSet or multi-replica Deployment duplicate data. Tail sampling needs all spans of a trace on one Collector, which is what the loadbalancing exporter is for. A sensible layout: a small agent per node that only collects, enriches and forwards, and a few gateway replicas that redact, sample and route with persistent queues.
Give the agent a memory limit that matches its buffer settings: a Collector without memory_limiter, or a Vector sink with a large in-memory buffer, can be OOM-killed under backpressure instead of slowing down.
OTLP and semantic conventions
OTLP is now the common language: Fluent Bit has native OTLP input and output, Vector has an opentelemetry source and sink, and the Collector is built around it. The latest releases are OTLP protobuf 1.11.1 (29 September 2026), specification 1.61.0 and semantic conventions 1.44.0 (5 August 2026).
Speaking OTLP is not the same as agreeing on field names. Semantic conventions define attributes such as service.name, k8s.pod.name and deployment.environment.name (the older deployment.environment is deprecated). If one cluster’s agent writes kubernetes.pod_name and another’s writes k8s.pod.name, queries break across clusters. Pick one scheme, set it at the agent, and test it with a sample record.
Resource use and benchmarks: whose run is it?
Most numbers come from the projects themselves. Fluent Bit’s release notes report its own gains; Vector’s README says “up to 10x faster than every alternative”, the Vector team’s own claim. A CNCF case study (12 November 2025) describes OpenAI cutting Fluent Bit CPU use by 50%: one user’s environment, not a comparison.
In my drill, only as a sense of scale, the highest sampled resident memory at 100 lines a second was about 19 to 21 MB for Fluent Bit, 52 to 61 MB for Vector and 221 to 226 MB for otelcol-contrib; the binaries were 76 MB, 156 MB and 407 MB. A slim ocb build would be smaller, but I did not measure one. A fair test uses your real lines, parsing rules and backend, and measures CPU per lakh lines, memory during a 30-minute outage, and loss and duplicates after a node drain.
Security advisories
All IDs below were confirmed as PUBLISHED through the CVE Services API on 8 October 2026.
Project
CVE
What happened
Fixed in
Fluent Bit
CVE-2025-12969
in_forward did not enforce security.users authentication in some configurations
4.0.13, 4.1.1 (per the Fluent Bit advisory)
Fluent Bit
CVE-2025-12972
out_file built file names from untrusted tags, allowing path traversal
4.0.13, 4.1.1
Fluent Bit
CVE-2025-12977
HTTP, Splunk and Elasticsearch inputs accepted unsanitised tag_key values
4.0.13, 4.1.1
Fluent Bit
CVE-2026-61674
Stack overflow in out_forward Secure Forward handshake with a malicious server
5.0.8
Fluentd
CVE-2026-44024
${tag} path traversal in out_file and similar outputs
1.19.3 (1.19.4 completes the fix)
Fluentd
CVE-2026-44025
in_monitor_agent exposed internal variables that may contain credentials
1.19.3
Fluentd
CVE-2026-44160
Compressed payloads to in_http and in_forward could exhaust memory
1.19.3
Vector
CVE-2026-77619, CVE-2026-77620
logstash source on 0.0.0.0:5044 could be crashed by oversized or nested compressed frames
0.57.0
Vector
CVE-2026-77621
file sink path templates allowed traversal from event fields
0.57.0
OTel Collector contrib
CVE-2026-42602
azure_auth extension authentication bypass
0.151.0
OTel Collector contrib
CVE-2026-55701
github receiver ignored required headers on webhooks
0.151.0
Logstash
CVE-2026-33466
Archive path traversal through the GeoIP database downloader
8.19.14, 9.2.8, 9.3.3
Grafana Alloy
CVE-2026-75889
prometheus.operator.servicemonitors could read local files and send them as a bearer token
1.19.0
The patterns repeat: listeners that trust whoever connects, tags or event fields used to build paths or URLs, and decompression without limits. Bind inputs to the network you mean, require TLS and authentication on forward and OTLP ports, never build output paths from tags you did not set, and run the agent as non-root. The Fluentd 1.19.4 notes say the 1.19.3 fix for CVE-2026-44024 was incomplete: take the latest patch, not the first fixed one.
Upgrade notes
- Fluent Bit 4.x to 5.1. In 5.0,
fluentbit_hot_reloaded_timesbecame a counter and HTTP input settings moved tohttp_server.*(old names still accepted). In 5.1, Debian and Ubuntu package upgrades restart a running service. Move from.confto YAML before the end-of-2026 deprecation. - Vector. Read every minor version’s upgrade guide; a jump from 0.50 to 0.59 means nine of them.
- OpenTelemetry Collector. Releases come about every two weeks. Read breaking changes for the components you use, rename deprecated aliases, and keep core and contrib versions in step in
ocbbuilds. - Fluentd. Stay on the latest 1.19.x or
fluent-packageLTS, and audit gems:fluent-plugin-s3(CVE-2026-44162) andfluent-plugin-opentelemetry(CVE-2026-44163) had advisories in September 2026. - Grafana Agent. Anything still running it is unpatched; move it to Alloy.
A decision guide
Your situation
Reasonable first choice
Kubernetes node agent, mostly logs, tight memory
Fluent Bit 5.1 with YAML, filesystem buffering and no_limits
Apps moving to OpenTelemetry SDKs; traces, metrics and logs together
OpenTelemetry Collector, slim build as agent plus gateway
Heavy parsing, redaction and routing with tested transforms
Vector as an aggregator with end-to-end acknowledgements
Grafana stack, many Prometheus scrape jobs
Grafana Alloy
Large Fluentd estate with custom Ruby plugins
Keep Fluentd patched and move node collection to Fluent Bit first
Elastic-centred platform
Elastic Agent or Logstash with persistent queues
Syslog from network and security devices
AxoSyslog or Fluent Bit’s syslog input, behind a gateway
For the Pune team, I would not rip everything out at once. Replace the Fluentd DaemonSet with Fluent Bit 5.1 in YAML, with filesystem buffering on a hostPath, retry_limit: no_limits and the dead letter queue, forwarding over OTLP to two or three Collector gateway replicas that redact with OTTL and keep a persistent queue sized for an hour of traffic. The apps’ SDKs send traces to the same gateways. Vector is a good gateway too if the team prefers VRL and acknowledgements; what matters more is writing down which hop owns durability, and testing it.
A practical checklist
- Write down the loss budget: how long an outage must be survived, and whether duplicates are acceptable.
- Check every retry default in the agent you run. Fluent Bit’s is 1; the Collector’s is 5 minutes.
- Put buffers on disk that survives a pod restart, and size it in hours of traffic.
- Persist read positions: Fluent Bit
db, Collectorfile_logstorage, Vectordata_dir. - Prefer acknowledgement or backpressure over silent drop wherever the source can wait.
- Alert on drops, not only on errors: retry exhaustion, queue full, dropped chunks, enqueue failures.
- Standardise attribute names with semantic conventions at the agent.
- Build a slim Collector with
ocbor use the k8s distribution instead of contrib in production. - Expose listeners deliberately, with TLS and authentication, and patch monthly.
- Rehearse an outage: stop the backend for 30 minutes and drain a node, then count what arrived.
Common mistakes
- Trusting Fluent Bit’s defaults for an outage. One retry and the chunk is discarded.
- Running the Collector’s
file_logwithout storage. A restart skips everything written while it was down. - Using
emptyDirfor buffers. It disappears with the pod; mount the buffer from the node. - Shipping
otelcol-contribto every node. It is a 407 MB binary of components you mostly do not use, each one attack surface.
Drill: losing logs on purpose on one VM
On 8 October 2026, from about 6:30 PM to 6:40 PM IST, on a shared Linux VM with 8 vCPUs (Intel Xeon) and 15.6 GiB of RAM, without root or containers, I ran each agent on 127.0.0.1 only and stopped every process afterwards.
- Fluent Bit 5.1.3 from the official Debian trixie package (SHA-256 matched the repository
Packagesindex, whoseInReleasefile had a good signature from the Fluent Bit release key), unpacked into my own directory withlibyamlfrom Debian. - Vector 0.59.0
x86_64-unknown-linux-musl(SHA-256 matchedSHA256SUMS). otelcol-contrib0.162.0 for linux_amd64 (SHA-256 matched the published.sha256file).
Method: a script wrote 6,000 numbered lines at 100 lines a second to a file; each agent tailed it and posted to a small Python HTTP receiver that recorded sequence numbers. The receiver was down from second 10 to second 40. In crash runs, the agent got SIGKILL at second 25 and was restarted at second 30, with the receiver still down.
Agent and settings
Crash?
Unique lines received of 6,000
Duplicates
Lost
Fluent Bit, defaults (memory, retry_limit 1)
No
3,800
0
2,200 (lines 973 to 3172); 22 “cannot be retried” log lines
Fluent Bit, memory, retry_limit: no_limits
No
6,000
0
0
Fluent Bit, filesystem buffering, db, no_limits
Yes
6,000
0
0
Vector, defaults
No
6,000
0
0
Vector, defaults
Yes
4,508
0
1,492 (lines 916 to 2407)
Vector, end-to-end acknowledgements, memory buffer
Yes
6,000
0
0
Vector, disk buffer (minimum size), no acknowledgements, two runs
Yes
6,000
89 each run
0
Collector, defaults
No
6,000
0
0
Collector, defaults
Yes
3,940
0
2,060 (lines 972 to 3031)
Collector, file_storage for file_log and sending_queue, three runs
Yes
5,980
0
20 each run (around lines 2475 to 2495)
Reading the results: Fluent Bit lost the chunks whose single retry fell inside the outage, and kept those whose retry came after the receiver returned. Vector’s loss is consistent with its file checkpoint having moved past lines that were still in memory. The Collector with defaults lost its in-memory queue at the kill plus the lines written while it was down, since file_log restarted at the end of the file. The 20 lines lost with storage sat just before the kill each time, consistent with the README’s warning about drops between components. In a separate test, Fluent Bit with retry_limit: 2 and storage.keep.rejected: on put the 50-line chunk it could not deliver into storage/rejected/ instead of dropping it.
Honest limits: one VM, one file, short lines, 100 lines a second, a 30-second outage and one kill per run. I did not test Kubernetes, log rotation, multiline parsing, TLS, real backends, node drains or high throughput. Memory figures are the highest sampled VmRSS during these runs, not a benchmark. Each setting was run once unless marked otherwise. I did not test Fluentd, Logstash, Alloy or AxoSyslog.
What to unlearn and re-learn
- Unlearn “the agent retries, so logs are safe”. Re-learn each agent’s retry limit and what happens when it is reached: discard, drop, or a dead letter queue.
- Unlearn “a disk buffer means no loss and no duplicates”. Re-learn that read positions, buffers and acknowledgements are three separate things, and you need all three to line up.
- Unlearn “OpenTelemetry means one stable thing”. Re-learn that each Collector component has its own stability level per signal, and that names still change between releases.
- Unlearn “this project will always be run the way it is today”. Re-learn who maintains each agent: a CNCF community, a single vendor, or a vendor that was just acquired.
Revisit your log pipeline before you ship more logs
The Pune team does not need the fastest agent; it needs one whose failure behaviour it understands. Learn where each agent keeps data when the backend says no, and for how long. Unlearn defaults inherited from a Helm chart written years ago. Re-learn buffering, read positions and acknowledgements from the current docs and release notes. Practise with a stopped backend and a killed agent, as I did on one VM, and count the lines. Then apply the result: a durable hop you can name, a loss budget you have tested, and a patch calendar for the agent on every node. Revisit the choice when your signals, your backends or the projects’ owners change.
Sources
- Fluent Bit repository, 5.1.3 release, 5.1.0 release, 5.0.10 release, 5.0.8 release, 5.0.0 release, 4.0.0 release, LICENSE
- Fluent Bit v5.1.3 notes, v5.1.0 notes, v5.0.0 notes, v4.0.0 notes, v3.0.0 notes, security fixes in v4.1 and v4.0
- Fluent Bit docs: history, Fluentd and Fluent Bit, buffering and storage, backpressure, scheduling and retries, dead letter queue, configuration formats, upgrade notes, Debian package repository
- CNCF: Fluentd project page, CNCF: Fluentd to Fluent Bit migration guide, CNCF: OpenAI and Fluent Bit case study, Chronosphere: Calyptia acquisition, Palo Alto Networks: agreement to acquire Chronosphere, Palo Alto Networks: acquisition completed
- Vector repository, README, LICENSE, 0.59.0 release, 0.58.0 release, 0.57.0 release, 0.59.0 release notes, vector.dev
- Vector docs: buffering model, end-to-end acknowledgements, VRL reference, http sink reference, Datadog: acquisition of Timber Technologies
- OpenTelemetry Collector repository, v0.162.0 release, Collector changelog, versioning and stability, component stability levels, exporter helper README, memory limiter processor
- Collector contrib repository, contrib v0.162.0 release, contrib changelog, file_log receiver README, OTTL README
- Collector releases repository v0.162.0, contrib distribution README, k8s distribution README
- OpenTelemetry docs: Collector distributions, resiliency, agent pattern, gateway pattern, Kubernetes components, transforming telemetry, semantic conventions, deprecated deployment attributes
- Semantic conventions v1.44.0, OTLP proto v1.11.1, specification v1.61.0
- CNCF: OpenTelemetry project page, CNCF: OpenTelemetry graduation announcement, OpenTelemetry blog: graduation
- Fluentd repository, v1.19.4 release, v1.0.0 release, Fluentd blog, buffer section docs
- Logstash repository, Logstash 9.5.5 release, Logstash LICENSE, persistent queues, ESA-2026-29 security update, Beats 9.5.5 release
- Grafana Alloy repository, Alloy 1.20.1 release, Alloy LICENSE, Alloy introduction, Alloy configuration syntax, Grafana blog: introducing Alloy, Grafana Agent docs (end of life notice), Grafana advisory CVE-2026-75889
- AxoSyslog repository, AxoSyslog 4.29.0 release, AxoSyslog COPYING
- CVE-2025-12969, CVE-2025-12972, CVE-2025-12977, CVE-2026-61674, CVE-2026-44024, CVE-2026-44025, CVE-2026-44160, CVE-2026-44162, CVE-2026-44163, CVE-2026-77619, CVE-2026-77620, CVE-2026-77621, CVE-2026-42602, CVE-2026-55701, CVE-2026-33466, CVE-2026-75889
- CVE Services API, Vector advisory GHSA-rrfg-9487-mhp6
Releted Posts
Argo CD and Flux: revisit your GitOps controller before you scale or migrate
A platform team in Pune looks after 14 Kubernetes clusters. Most of them are managed by one central Argo CD instance that the team set up in 2021.
Read moreIngress NGINX retirement: revisit your Kubernetes ingress before you migrate
A platform team plans its Kubernetes 1.37 upgrade for the next sprint. The cluster checklist is green, except one line nobody owns: the ingress-nginx Helm chart, pinned at 4.
Read moreService mesh in 2026: revisit Istio ambient, Linkerd and Cilium before you add sidecars
A logistics company in Bengaluru runs a production Kubernetes cluster with 14 nodes and about 420 pods. Three requests landed in one sprint.
Read more