Prometheus long-term storage: revisit Thanos, Mimir and VictoriaMetrics before you scale

A platform team runs two Prometheus servers per cluster as an HA pair, with 15 days of local retention. In one quarter, three requests arrive: the SRE lead wants a year of history for capacity planning, a product team adds a customer_id label to a request metric, and the auditors ask what happens to metrics when a disk dies. One person says “add Thanos”, another says “Mimir”, a third has read that VictoriaMetrics is cheaper. All three are reasonable answers to different questions.

This post covers Prometheus local storage limits, the history and state of Cortex, Thanos, Grafana Mimir and VictoriaMetrics, architectures, licences, remote write 2.0, native histograms, PromQL differences, HA deduplication, downsampling and limits, with a decision guide and a checklist. Facts come from the projects’ repositories, release notes, changelogs, licence files and docs, and CNCF pages, checked on 7 October 2026. Drill results are from my own run. The earlier post on MinIO and its alternatives covered object storage; this one is about the metrics layer on top of it.

The short version

  • Prometheus local storage is a single-node database: 15 days by default, not clustered, not replicated. Long-term storage means a second system, fed either by shipping TSDB blocks to object storage or by remote write.
  • Current releases: Prometheus 3.15.0 (24 September 2026; 3.13 is the LTS line), Thanos 0.42.4 (30 July 2026), Grafana Mimir 3.2.1 (10 September 2026), VictoriaMetrics 1.153.0 (28 September 2026) and Cortex 1.21.1 (June 2026).
  • Thanos, Cortex and VictoriaMetrics community edition are Apache 2.0; Mimir is AGPLv3. Downsampling in VictoriaMetrics is an enterprise feature, and Mimir does not downsample at all.
  • Remote write 2.0 is still an experimental spec. In my drill, Prometheus 3.15.0 sending 2.0 succeeded against Mimir 3.2.1 and failed against VictoriaMetrics 1.153.0.
  • Each system removes HA duplicates differently. When I stopped the elected replica, Mimir showed a 41 second gap; Thanos and VictoriaMetrics showed none.
  • The same counter gave an increase() of 543.27 on Prometheus, Thanos and Mimir and 550 on VictoriaMetrics, because MetricsQL does not extrapolate.

Where things stand on 7 October 2026

Thanos

Grafana Mimir

VictoriaMetrics

Cortex

Latest release

0.42.4, 30 July 2026

3.2.1, 10 September 2026; 3.1.7 patch, 7 October 2026

1.153.0, 28 September 2026; enterprise LTS lines 1.148.x and 1.136.x

1.21.1, June 2026; 1.22.0-rc.3, 7 October 2026

Licence

Apache 2.0

AGPLv3

Apache 2.0; enterprise build needs a licence key

Apache 2.0

Home

CNCF incubating

Grafana Labs

VictoriaMetrics company

CNCF incubating

How data arrives

Sidecar uploads blocks, or Receive takes remote write

Remote write, OTLP

Remote write, scraping, many push formats

Remote write

Where data lives

Object storage

Object storage

Local disks

Object storage

Query language

PromQL

PromQL, Mimir Query Engine by default

MetricsQL

PromQL

Downsampling

Yes, in the compactor

No

Enterprise only

No

The timeline, from the repositories, release notes and CNCF pages:

Date

Event

22 June 2016

First Cortex commit; started at Weaveworks by Tom Wilkie and Julius Volz

January 2018

Fabian Reinartz and Bartłomiej Płotka of Improbable present Thanos at the London Prometheus meetup

20 September 2018

Cortex joins the CNCF sandbox

30 September 2018

VictoriaMetrics repository created; first open-source release in May 2019

14 July 2019

Thanos joins the CNCF

August 2020

Thanos (19 August) and Cortex (20 August) move to incubating

March 2022

Grafana Labs announces Mimir, a fork of Cortex, under AGPLv3

14 November 2024

Prometheus 3.0.0; agent mode becomes stable

31 October 2025

Mimir 3.0.0: Kafka-based ingest storage, Mimir Query Engine as default

28 November 2025

Prometheus 3.8.0 makes native histograms stable, opt-in per scrape config

1 July 2026

Prometheus 3.13 LTS, supported until 31 July 2027

24 September 2026

Prometheus 3.15.0 stabilises the XOR2 float chunk encoding

What Prometheus local storage promises

The storage docs are direct: local storage “is not clustered or replicated”, so it is “not arbitrarily scalable or durable in the face of drive or node outages”. They also say years of data can be kept locally with proper architecture. So the usual reasons for a second system are:

  • Durability. One disk failure loses that server’s history.
  • A global view. Ten clusters means ten data sources in Grafana.
  • Long range queries. A year of raw 15 second samples is a lot to read.
  • Cardinality. Active series live in memory, so one bad label can exhaust a server.

A few facts to know first. Retention defaults to 15 days, and 3.15.0 marks the retention flags deprecated in favour of storage.tsdb.retention in the config file. Compaction builds blocks up to 10% of retention or 31 days, whichever is smaller. The docs suggest size-based retention of at most 80 to 85% of the disk, say NFS is not supported, and quote 1 to 2 bytes per sample for planning. Agent mode (--agent, stable since 3.0) scrapes and remote writes without local query storage, which suits setups where all querying happens elsewhere.

Two ways to get data out of Prometheus

Ship blocks. The Thanos sidecar uploads each completed two-hour block to object storage and serves recent data over its gRPC StoreAPI. Local compaction must be off: minimum and maximum block duration set equal, 2h in the docs.

Push samples. Prometheus remote writes every sample to Thanos Receive, Mimir, Cortex or VictoriaMetrics, and can then run with short retention or as an agent. The queue can fall behind, and the receiver’s limits now decide what is accepted.

Thanos sidecar

Thanos Receive

Mimir and Cortex

VictoriaMetrics single

VictoriaMetrics cluster

Write path

Block upload every 2h

Remote write

Remote write

Remote write

Remote write to vminsert

Needs object storage

Yes

Yes

Yes

No

No

Main components

Sidecar, Query, Store, Compactor; optional Query Frontend, Ruler

Receive plus the same read side

Distributor, ingester, querier, query-frontend, query-scheduler, store-gateway, compactor, ruler; Kafka with ingest storage

One binary

vminsert, vmselect, vmstorage

The projects, one by one

Cortex: where push-based storage started

Cortex began in 2016 as a horizontally scalable, multi-tenant Prometheus service and is still a CNCF incubating project under Apache 2.0. Grafana Labs’ Mimir announcement says its employees made about 87% of Cortex commits from 2019 to 2021, so the 2022 fork moved most of the effort. Cortex continues: 1.21.1 shipped in June 2026, 1.22.0 is in release candidates, and its changelog keeps extending Parquet-based storage.

Thanos: Prometheus plus object storage

Thanos came from Improbable, which needed a global view and long retention across many Kubernetes clusters. It leaves Prometheus as it is and adds components around it. The querier fans out to sidecars, store gateways and receivers, and removes HA duplicates at query time using a replica label. The compactor merges blocks, applies retention and downsamples: 5 minute resolution for blocks older than 40 hours, 1 hour for blocks older than 10 days, with separate retention flags for raw, 5m and 1h data. Native histogram downsampling arrived in 0.38.0.

The querier uses the Prometheus engine by default; a Thanos engine and a distributed mode are optional. The docs recommend Receive for setups that can only push, such as air-gapped or egress-only networks.

Grafana Mimir: Cortex rebuilt by Grafana Labs

Mimir was announced in March 2022 as a Cortex fork under AGPLv3, bringing formerly commercial features such as the split compactor and query sharding into open source. It runs as one binary (-target=all) or as microservices, which the docs recommend for production.

Mimir 3.0 added ingest storage, with Kafka between the write and read paths; the classic architecture is still supported. It also made the Mimir Query Engine the default, described as “fully PromQL-compatible”, with -querier.query-engine=prometheus as the way back. Mimir does not downsample, and its migration guide says Thanos downsampled blocks must not be uploaded because raw and downsampled samples would be merged at query time.

VictoriaMetrics: a different storage engine

VictoriaMetrics was written by Aliaksandr Valialkin and has been open source under Apache 2.0 since May 2019. It is not a CNCF project and does not use the Prometheus TSDB format. It accepts remote write, scraping and many other protocols.

The single node is one binary on local disk, with a default retention of one month. The cluster is shared nothing: vminsert spreads series across vmstorage nodes by consistent hashing, and vmselect queries them all. Tenants exist only in the cluster, by accountID. Neither needs object storage; vmbackup copies snapshots to it. The enterprise edition adds downsampling, multiple retentions and LTS releases, and community users are told to stay on the latest release.

Licences: read them before the architecture review

  • Apache 2.0 (Thanos, Cortex, VictoriaMetrics community): permissive.
  • AGPLv3 (Mimir): if you modify it and users interact with it over a network, you must offer them the modified source. Running it unmodified inside your company is normal use, but many legal teams still want a review. Plan that time.
  • Open core (VictoriaMetrics): downsampling and per-dataset retention need an enterprise licence. Decide early whether you will pay or design without them.

Remote write 2.0 and native histograms

The remote write 2.0 spec is still marked experimental (2.0-rc.5). It adds a symbol table, metadata, start timestamps, native histograms and response headers such as X-Prometheus-Remote-Write-Samples-Written. Prometheus uses it only with protobuf_message: io.prometheus.write.v2.Request; the default is still 1.0. Mimir’s 3.0 notes list 2.0 support as experimental.

The headers matter: the spec tells senders to treat a 2xx without “Written” headers as a receiver that does not really understand 2.0. In my drill, VictoriaMetrics 1.153.0 answered exactly that way, so Prometheus dropped every 2.0 batch as a non-recoverable failure. Mimir accepted all of them. Keep 1.0 unless your receiver documents 2.0 support.

Native histograms have been stable in Prometheus since 3.8.0, opt-in with scrape_native_histograms: true, and 1.0 remote write queues also need send_native_histograms: true. Thanos Receive has ingested them by default since 0.41.0, and Mimir 3.2.1 does by default although its flag is still marked experimental. VictoriaMetrics converts them into its own bucket series with a vmrange label: in my drill, 200 native histogram series became more than 7,400 series.

PromQL and MetricsQL are not the same

The MetricsQL docs list intentional differences. Two matter most:

  • rate() and increase() use the last sample before the window and are not extrapolated. Prometheus extrapolates to the window edges, so increase() over an integer counter can return a fraction.
  • A range shorter than the scrape interval still returns a value in MetricsQL. In Prometheus, rate(x[5s]) with a 5 second scrape returns nothing.

Neither is wrong. The risk is migration: alert thresholds tuned on Prometheus numbers, SLO reports compared across the switch, recording rules moved between engines. In my drill, Prometheus, Thanos with dedup on and Mimir returned identical values for every query I compared.

Deduplication of HA pairs

Most teams run two identical Prometheus servers that differ only in a replica external label.

Where dedup happens

Copies stored

When one replica dies

Thanos

At query time, --query.replica-label; penalty algorithm by default

Two, unless the compactor deduplicates

The other replica’s data is already there

Mimir and Cortex

At the distributor; the HA tracker elects one replica per cluster

One

The other replica’s samples are dropped until failover; default timeout 30 seconds

VictoriaMetrics

One sample per -dedup.minScrapeInterval; labels must match, so drop the replica label

One, after background merges

The other replica fills in

Mimir saves storage and keeps queries simple, but the failover window is a real gap. In my drill, the last sample from the stopped replica was at 10:41:02 IST and the first from the new one at 10:41:43.

Downsampling, retention and limits

Thanos downsampling speeds up long range queries but saves little space unless raw retention is shorter, because the compactor keeps raw, 5m and 1h copies by their own retention flags. Mimir does not downsample; in a community Q&A, its team said object storage is about 5% of Mimir’s total cost in their experience. VictoriaMetrics offers it in enterprise. Retention is per resolution in Thanos, per tenant in Mimir (-compactor.blocks-retention-period, off by default) and per instance in VictoriaMetrics community.

Mimir and Cortex are multi-tenant through the X-Scope-OrgID header. Mimir 3.2.1 defaults to 150,000 in-memory series per tenant, 10,000 samples per second and a burst of 200,000, which a busy cluster will hit on day one. VictoriaMetrics has a cardinality limiter (-storage.maxHourlySeries, -storage.maxDailySeries) that drops new series above the limit. In Prometheus itself, sample_limit and label_limit per scrape job are the cheapest first defence.

Operations: count the moving parts

Production Thanos means sidecars on every Prometheus plus query, query frontend, store gateways, a compactor per bucket and often a ruler. Mimir in microservices mode means eight components, a key-value store for the rings, and Kafka if you choose ingest storage. VictoriaMetrics cluster has three services, the single node one. Fewer components is not zero work: VictoriaMetrics keeps data on local disks, so you size them, plan replicas and run backups, while Thanos and Mimir push durability to object storage that you run or rent.

A starting remote write block for one server of an HA pair:

global:
  scrape_interval: 15s
  external_labels:
    cluster: prod-mumbai-1
    replica: a                     # b on the second server
remote_write:
  - url: http://mimir-gateway/api/v1/push
    headers:
      X-Scope-OrgID: platform
    send_native_histograms: true   # needed with the 1.0 message
  - url: http://victoriametrics:8428/api/v1/write
    # keep the 1.0 message here; drop "replica" on the receiver
    # and set -dedup.minScrapeInterval=15s

Performance: claims and what I measured

I found no recent independent benchmark comparing Thanos, Mimir and VictoriaMetrics on the same hardware and workload. The projects’ own figures:

  • Grafana Labs’ own claims: Mimir is “up to 40x faster than Cortex” (2022 announcement), and MQE cuts peak memory “by up to 92%” against the Prometheus engine (3.0 notes).
  • VictoriaMetrics’ own claims: “up to 7x less RAM” than Prometheus, Thanos or Cortex with millions of series, and “up to 7x less storage space”, citing its own benchmarks.
  • Prometheus docs: “an average of only 1-2 bytes per sample”.

Read these as project claims. My numbers come from one small machine and about 18 minutes of data.

A decision guide

Your situation

Reasonable first choice

Prometheus works; need history and a global view

Thanos with sidecars and object storage

Year-long dashboards that stay fast

Thanos with downsampling, or VictoriaMetrics enterprise

Many teams or customers with strict limits

Mimir, or Cortex if you already run it

Small team, one region, fewest components

VictoriaMetrics single node, with backups

Beyond one node, no object storage

VictoriaMetrics cluster

Push-only or air-gapped sites

Remote write to Mimir, Thanos Receive or VictoriaMetrics

AGPL not acceptable

Thanos, Cortex or VictoriaMetrics

Alerts must give the same numbers after the move

Thanos or Mimir, or a planned rule review for MetricsQL

A practical checklist

  1. Write down the real need: retention, durability, global view, tenants or cost.
  2. Measure the load: active series, samples per second, churn and the top metrics by series.
  3. Fix cardinality at the source with sample_limit, label_limit and relabelling.
  4. Choose how HA pairs are deduplicated and test a replica failure.
  5. Check licences and enterprise-only features against your needs.
  6. Keep remote write 1.0 unless the receiver documents 2.0, and check the “Written” headers when testing.
  7. Enable native histograms in one job first and see what the receiver stores.
  8. Compare your top alerts and SLO queries on both engines before switching.
  9. Set receiver limits deliberately: series per tenant, rate and burst.
  10. Rehearse a restore from the bucket or backup.

Common mistakes

  • Using long-term storage to fix cardinality. The bad label moves somewhere bigger and costlier.
  • Assuming dedup is free. Thanos stores both replicas unless the compactor deduplicates; in my drill, the bucket held both.
  • Keeping Mimir’s default limits on a production tenant.
  • Expecting Mimir to downsample, or uploading Thanos downsampled blocks into it.
  • Comparing increase() across engines without knowing about extrapolation.
  • Trusting a 2xx on remote write 2.0 without the response headers.

Drill: one HA pair, three stores, five surprises

I ran this on 7 October 2026, about 10:24 to 10:45 AM IST, on a shared VM container: 8 vCPUs (Intel Xeon), 15 GiB RAM with about 6.4 GiB available, Debian 13, Linux 6.12. Official release binaries: two Prometheus 3.15.0 servers (A and B) as an HA pair, Thanos 0.42.4 (two sidecars, store and query on a local filesystem bucket), Mimir 3.2.1 monolithic with filesystem storage and the HA tracker on, and VictoriaMetrics 1.153.0 single node with -dedup.minScrapeInterval=5s and the replica label dropped. Load came from avalanche 0.7.0: 20 gauge and 20 counter metrics with 100 series each, plus 2 native histogram metrics with 100 series each, changing every 5 seconds and scraped every 5 seconds, so 4,205 series per server. Both servers wrote to Mimir and VictoriaMetrics. I set the block length to 5 minutes so the sidecars would upload during the drill. Everything ran on localhost and was stopped afterwards.

Ingestion and deduplication

No remote write queue on the HA pair reported a failed sample. By 10:40:48, Mimir had received 1,485,255 samples, dropped 740,700 from the non-elected replica and kept 744,555. After a forced merge at the end, VictoriaMetrics had received 4,272,803 rows (native histograms expanded into buckets), removed 1,960,136 duplicates and stored 2,276,243. The Thanos bucket held 1,278,320 samples in 7 blocks from both replicas.

Same data, same moment, different answers

Queries at 10:40:00 IST; c is one counter metric with 100 series:

Query

Prometheus A

Thanos, dedup on

Thanos, dedup off

VictoriaMetrics

Mimir

Counter series, all metrics

2,000

2,000

4,000

2,000

2,000

Native histogram series

200

200

400

7,475

200

sum(rate(c[1m]))

989

989

1,978

991.05

989

sum(increase(c[1m]))

59,340

59,340

118,680

59,463

59,340

increase(), one series

543.27

543.27

2 series

550

543.27

count(rate(c[5s]))

empty

100

empty

100

empty

p90 of a native histogram

88.435

88.435

88.435

88.750, from vmrange buckets

88.435

Mimir followed replica A, so it matched Prometheus A exactly. VictoriaMetrics keeps one sample per 5 second window from either replica and does not extrapolate, so its values differed slightly and its single-series increase() was an integer. With dedup off, Thanos doubled every sum, which is what a dashboard shows when the replica label is not configured. One surprise: with dedup on, Thanos returned 100 results for the 5 second window, where each replica alone returned none, apparently because the merged stream had samples from both replicas inside it. I did not dig further.

Stopping the elected replica

I stopped Prometheus A at 10:41:07. Mimir’s HA tracker elected B at 10:41:40. Samples of one gauge series in the next 60 seconds, 12 expected: Prometheus B 12, Thanos with dedup 12, VictoriaMetrics 12, Mimir 5. The failover timeout can be tuned, but electing one replica always costs some gap.

Remote write 2.0

A third Prometheus 3.15.0 sent the 2.0 message to both receivers for about a minute. Mimir accepted 32,000 of 32,000 samples. VictoriaMetrics answered 2xx without “Written” headers, and Prometheus logged each batch as a non-recoverable failure; nothing was stored. That queue also sent 0 histograms although the server appended 2,200 histogram samples, and Mimir held only float series for it. Separately, promtool push metrics with the 2.0 message got a 400 from Mimir while 1.0 worked. I did not find the causes in the time box; treat these as things to test, not as known bugs.

Disk and memory, at toy scale

Store

Bytes on disk

Samples

Bytes per sample

Prometheus A snapshot, four blocks of up to 5 minutes

4,158,206

819,975

5.07; chunks alone 2.81

Thanos bucket, both replicas

6,740,348

1,278,320

5.27

Mimir, one flushed block

2,706,643

873,975

3.10; chunks alone 2.52

VictoriaMetrics after a forced merge

3,349,081

2,276,243 rows

Not comparable: 1,369,306 rows were expanded histograms

All are well above the 1 to 2 bytes per sample in the docs, most likely because a few minutes of 4,205 series is mostly index and native histograms rather than long compressed chunks. Resident memory at 10:40:48: about 175 to 180 MB per Prometheus, 308 MB for the four Thanos processes together, 269 MB for Mimir and 131 MB for VictoriaMetrics.

Honest limits: one machine, localhost, about 18 minutes of data, 4,205 synthetic series, single runs, filesystem buckets instead of an object store. Mimir was restarted once early on to fix its query-frontend address. I did not test Thanos Receive, the compactor, downsampling, VictoriaMetrics cluster, Mimir ingest storage or Cortex. The byte counts are not a storage benchmark.

What to unlearn and re-learn

  • Unlearn “long-term storage is just more retention”. Re-learn it as a choice about durability, global view, tenants and cost.
  • Unlearn “they all speak PromQL, so the numbers match”. Re-learn that MetricsQL defines rate() and increase() differently, and test your alerts.
  • Unlearn “HA pairs are handled”. Re-learn where each system deduplicates, what it stores and what a failover costs.
  • Unlearn “remote write 2.0 is the default now”. Re-learn that it is opt-in and experimental, and that a 2xx is not proof.

Revisit your metrics storage before you scale

Prometheus is still the right collector, and each long-term store answers a different question: Thanos keeps Prometheus as it is, Mimir runs metrics as a multi-tenant service, and VictoriaMetrics brings its own engine with fewer moving parts. Learn what your series and samples really are, unlearn the idea that every PromQL-shaped API gives the same numbers, re-learn deduplication and downsampling for the tool you pick, practise the drill with your own HA pair and top alerts, and apply the results in your limits, retention policy and runbooks. Do it before retention or cardinality forces the choice.

Sources

comments powered by Disqus

Releted Posts

PgBouncer and its alternatives: revisit your PostgreSQL connection pooler

A team runs about forty services on Kubernetes against one PostgreSQL primary. During a sale, the autoscaler adds pods, each opens its own pool of ten connections, and PostgreSQL starts refusing logins with “remaining connection slots are reserved”.

Read more

Vault and OpenBao: revisit your secrets manager before you renew or migrate

A platform team runs a three-node Vault Community cluster for about sixty services. In one week, a product group asks for its own isolated tenant, which in Vault means namespaces, an Enterprise feature, and finance asks whether the Enterprise quote is really needed.

Read more

Redis, Valkey, and Dragonfly clients: revisit your connection settings before the next failover

An order service runs on Valkey with Sentinel. At 2 AM the team runs a planned failover before patching the primary.

Read more