Vector search in 2026: revisit pgvector, Qdrant, Milvus, Weaviate before you add a vector database
An e-commerce company in Pune runs customer support on PostgreSQL. The database holds about 2 million help articles, product Q&A threads and resolved tickets, in English and a fair amount of Hinglish. Before the festive season, the product team wants semantic search for support agents (“refund never came” should find articles that say “reversal pending”) and a RAG assistant on the help page. After chunking, that is about 6 million chunks with 768-dimension embeddings. In the design review, one engineer proposes Qdrant, another has a Milvus slide deck from a meetup, and the DBA asks why they cannot simply install pgvector on the database they already back up, monitor and patch.
That question deserves a serious answer. A separate vector database brings its own storage, backups, auth, upgrade calendar and security advisories. Sometimes that is worth it; often it is not. This post compares pgvector (with pgvectorscale and VectorChord), Qdrant, Milvus and Weaviate on history, index types, memory, filtering, hybrid search, licences, operations and security, with short notes on OpenSearch, Elasticsearch, Redis, Valkey, Chroma and LanceDB. Facts come from repositories, release pages, licence files, official docs, papers and CVE records, checked on 8 October 2026; the drill is my own run.
The short version
- pgvector 0.8.7 (1 October 2026) is current. Three pgvector CVEs were published in 2026, all reachable by a database user through index builds.
- For a few million vectors with ordinary filters, PostgreSQL with pgvector is usually enough. Move out when vectors outgrow one machine’s memory, when search load hurts the OLTP workload, or when you need features PostgreSQL lacks.
- Filtering is where most teams get wrong answers. In my drill, a filtered HNSW query in pgvector with default settings returned 0.4 rows on average instead of 10. The fix is a setting, not a new database.
- Memory decides the architecture: 6 million float32 vectors of 768 dimensions are about 17.2 GiB before any index. Half precision halves that; 1-bit quantisation brings it near 0.54 GiB, at a recall cost you must measure.
- Weaviate’s LICENSE file changed in a patch release, 1.39.4 (11 September 2026): BSD-3-Clause outside a new
wldirectory, a commercial licence inside it. In 1.40.0 (7 October) that directory holds features such as namespaces and deduplicated backups. - Most published benchmarks are vendor runs. Shortlist with them, then test your own embeddings and filters.
Where things stand on 8 October 2026
Release dates are from GitHub release pages, in IST.
Project
Latest release (IST)
What it is
Licence (LICENSE file)
pgvector
0.8.7 on 1 October 2026
PostgreSQL extension; HNSW and IVFFlat
PostgreSQL License
pgvectorscale
0.9.1 on 4 September 2026
Extension on pgvector: StreamingDiskANN, binary quantisation
PostgreSQL License
VectorChord
1.1.1 on 28 February 2026
Extension on pgvector types: IVF with RaBitQ
AGPLv3 or Elastic License v2
Qdrant
1.19.2 on 5 October 2026
Vector database in Rust
Apache 2.0
Milvus
3.0.2 on 20 September; 2.6.25 on 29 September 2026
Distributed vector database, LF AI & Data graduated
Apache 2.0
Weaviate
1.40.0 and 1.39.10 on 7 October 2026
Vector database in Go with BM25 and hybrid search
BSD-3-Clause, plus commercial wl directory since 1.39.4
Chroma
1.5.9 on 5 May 2026
Developer-first vector store
Apache 2.0
LanceDB
0.40.0 on 7 October 2026
Embedded, on the Lance columnar format
Apache 2.0
valkey-search
1.2.1 on 8 July; 1.3.0-rc1 on 7 October 2026
Valkey module: vector, text, tag, numeric
BSD 3-Clause
Also current: Redis 8.10.2 (17 September 2026) with vector sets, OpenSearch k-NN 3.9.0.0 (22 September), Elasticsearch 9.5.5 (6 October) and FAISS 1.15.1 (17 September).
Date
Event
30 March 2016
HNSW paper by Malkov and Yashunin on arXiv
March 2017
Facebook releases Faiss (blog post of 29 March); its GPU paper was on arXiv from 28 February
December 2019
DiskANN at NeurIPS: a billion points on one machine with 64 GB RAM and an SSD
14 January 2021
Weaviate 1.0.0
10 March 2021
Milvus 1.0.0; it graduates in LF AI & Data on 23 June 2021
20 April 2021
pgvector 0.1.0
8 February 2023
Qdrant 1.0.0
28 August 2023
pgvector 0.5.0 adds HNSW
6 June 2024
pgvectorscale 0.2.0 from Timescale
30 October 2024
pgvector 0.8.0 adds iterative index scans
2 May 2025
Redis 8.0 with vector sets; valkey-search 1.0.0 follows on 28 May
29 July 2026
Milvus 3.0.0, the “lake-native” release
11 September 2026
Weaviate 1.39.4 adds the BSD plus commercial wl split
7 October 2026
Weaviate 1.40.0: namespaces and deduplicated backups need a licence key
Do you need a separate vector database at all?
Vector search finds the k stored embeddings nearest to a query embedding. The maths is not exotic; memory, filtering and index freshness are what make it hard at scale.
PostgreSQL with pgvector covers more ground than many teams expect. Vectors sit in the same rows as ticket text and tenant IDs, so filters are ordinary SQL, and the pgvector README confirms it uses the write-ahead log, so replication and point-in-time recovery just work. Existing backups, pooling and monitoring keep working; the earlier posts on PostgreSQL asynchronous I/O and PgBouncer and its alternatives cover those skills.
A dedicated engine earns its place when one of these is true:
Signal
Why PostgreSQL struggles
What a dedicated engine offers
Vectors plus index outgrow the RAM of your largest affordable primary
HNSW wants its working set in memory, and every replica copies it
Sharding, on-disk indexes, memory tiers
Search and re-indexing compete with order traffic
One buffer cache, one set of CPUs
Separate hardware and scaling
Heavy filtering on many fields
Post-filtering on approximate indexes needs care
Filter-aware graph search, payload indexes
If none applies today, start with pgvector and measure. Moving 6 million vectors later is a batch job, not a rewrite, if the application calls one small retrieval interface.
How approximate search works
Exact search compares the query with every vector: about 27 ms per query for 100,000 vectors in my drill, growing roughly linearly. Approximate nearest neighbour (ANN) indexes trade a little recall for a lot of speed. Recall@10 is the share of the true top 10 that the index returns.
HNSW
Hierarchical Navigable Small World graphs link each vector to a few neighbours on several layers, and a search walks greedily towards the query. Three knobs matter:
m, links per node: more links, better recall, more memory. pgvector and Qdrant default to 16; Weaviate’smaxConnectionsdefaults to 32.ef_construction, the build-time candidate list: pgvector 64, Qdrant 100, Weaviate 128. Higher means a better graph and a slower build.ef_search(Qdranthnsw_ef, Weaviateef), the query-time candidate list: the main recall versus latency dial, adjustable per query. pgvector defaults to 40; Weaviate picks a dynamic value.
HNSW gives excellent recall and latency but wants RAM, and deletes need vacuuming or compaction.
IVF and DiskANN
IVF indexes cluster vectors with k-means into lists and search only the nearest probes lists. Builds are lighter, but clusters are trained on existing data, so an IVF index built on a tiny table stays poor. pgvector’s README suggests rows / 1000 lists up to 1 million rows, sqrt(rows) above that, and sqrt(lists) probes. VectorChord builds on IVF.
DiskANN keeps most data on SSD and compressed vectors in memory; its NeurIPS 2019 paper reports, on the authors’ hardware, over 5,000 queries per second at 95% recall on a billion points with 64 GB of RAM. pgvectorscale and Milvus’s DISKANN index follow this idea.
Quantisation
Method
Bytes per 768-dim vector
Compression
float32
3,072
1x
Half precision (halfvec, FLOAT16)
1,536
2x
Scalar int8 (Qdrant scalar, Weaviate SQ, Elasticsearch int8_hnsw)
768
4x
4-bit (Qdrant TurboQuant, Weaviate RQ)
384
8x
Binary, 1 bit per dimension
96
32x
Product quantisation (trained codebooks)
Configurable
Up to about 64x
The rule that prevents most mistakes: compressed vectors find candidates, original vectors rank them. Rescoring defaults differ: Qdrant’s docs say it rescores by default only for binary and low-bit TurboQuant, not scalar int8, and my drill shows the cost.
Memory arithmetic before product choice
Do this before any vendor call. As a reference: 10 million vectors x 768 dimensions x 4 bytes = 30,720,000,000 bytes, about 28.6 GiB of raw floats before any index. For the Pune team’s 6 million chunks:
Representation
Bytes per vector
6 million vectors
pgvector vector (4 x dims + 8)
3,080
about 17.2 GiB
pgvector halfvec (2 x dims + 8)
1,544
about 8.6 GiB
int8 scalar
768
about 4.3 GiB
4-bit
384
about 2.1 GiB
1-bit binary
96
about 0.54 GiB
HNSW layer-0 links, m = 16 (32 links of 4 to 6 bytes)
128 to 192
about 0.7 to 1.1 GiB
In pgvector the HNSW index keeps its own copy of each vector beside the graph. In my drill the overhead beyond the vector was about 500 bytes per row for a float32 HNSW index and about 400 bytes for halfvec. Extrapolating, a float32 HNSW index on 6 million 768-dim vectors would be about 20 GiB on top of an 18 GiB table, and a halfvec expression index about 11 GiB. An estimate from a small run, but the conclusion holds: a halfvec index fits well in a 64 GB server; a float32 one fights everything else for cache.
Count replicas too. Three copies of a 20 GiB index is 60 GiB of RAM you pay for.
pgvector, pgvectorscale and VectorChord
pgvector 0.8.x is now a maintenance line: 0.8.2 to 0.8.7 (February to October 2026) are fixes, including possible HNSW index corruption during vacuum (0.8.3) and the security bugs below. Index limits are 2,000 dimensions for vector, 4,000 for halfvec and 64,000 for bit. The settings most teams need can be set per transaction:
-- Index half-precision copies; keep full precision in the table
CREATE INDEX CONCURRENTLY chunks_embedding_hnsw
ON chunks USING hnsw ((embedding::halfvec(768)) halfvec_cosine_ops)
WITH (m = 16, ef_construction = 64);
BEGIN;
SET LOCAL hnsw.ef_search = 100; -- default 40
SET LOCAL hnsw.iterative_scan = relaxed_order; -- keep scanning when filters drop rows
SELECT id, title
FROM chunks
WHERE tenant_id = 42 AND language = 'hi'
ORDER BY embedding::halfvec(768) <=> $1::halfvec(768)
LIMIT 10;
COMMIT;
The query must repeat the index expression (embedding::halfvec(768)) or the planner ignores the index. For builds, give maintenance_work_mem room for the graph without starving the server.
pgvectorscale, from Timescale (now Tiger Data), adds a StreamingDiskANN index with Statistical Binary Quantization and label-based filtering, in Rust, under the PostgreSQL License. Its 0.9.1 release is a security hardening fix: malformed vector data or a type pre-created in its schema could cause a crash, memory disclosure or out-of-bounds write. Its README claims 28x lower p95 latency and 16x higher throughput than Pinecone’s s1 index at 99% recall on 50 million 768-dim Cohere embeddings; that is the vendor’s own benchmark.
VectorChord uses pgvector’s types with its own vchordrq IVF index and RaBitQ quantisation; 1.1.0 added rabitq8 and rabitq4 types. Its README claims 100 million vectors indexed in 20 minutes, again the vendor’s own number. Its licence, AGPLv3 or Elastic License v2, is a real decision for anyone who may offer the database inside a hosted product.
Qdrant
Qdrant is a single Rust binary with REST and gRPC APIs, JSON payloads, payload indexes and a Raft-based cluster mode. 1.19.0 (5 August 2026) added 4-bit TurboQuant as a storage type for primary vectors and per-component memory settings (cold, cached, pinned). 1.19.2 fixed HNSW builds not using all assigned threads and added a startup warning when no API keys are set.
- Security is off until you turn it on. The docs say API keys, TLS and audit logging are on by default in Qdrant Cloud but must be configured on self-hosted deployments. My drill binary logged “TLS disabled” and answered every request without a key.
- Shards are fixed at creation.
shard_numberdefaults to the node count when the collection is created and cannot change without recreating it. - Snapshots restore only on the same or the next minor version, and aliases let you swap in a re-embedded collection.
Milvus
Milvus is the most distributed of the four: stateless proxies, streaming, query and data nodes, etcd for metadata, MinIO or S3 for data and indexes, and Woodpecker as the write-ahead log. Milvus 3.0.0 (29 July 2026) added External Collections that index data-lake files in place, online schema changes and Woodpecker as a standalone service. The 2.6 line is still patched.
Its index menu is the widest: IVF variants including IVF_RABITQ, HNSW with SQ, PQ and PRQ, DISKANN, SCANN and GPU indexes such as GPU_CAGRA. Valuable at hundreds of millions of vectors; at 6 million, mostly extra moving parts. Authentication is opt-in (authorizationEnabled: true), and the default root password is Milvus, so change it on day one.
Weaviate
Weaviate, in Go, is strong for RAG: built-in BM25F and hybrid search that fuses keyword and vector results with a tunable alpha. Index types are hnsw (default), flat, dynamic (experimental) and hfresh; compression options are BQ, PQ, SQ and rotational quantisation (RQ). Since 1.33 a DEFAULT_QUANTIZATION variable can compress new collections, and the docs warn that quantisation cannot be switched off once set.
The licence change is the news. Up to 1.39.3 the LICENSE file was plain BSD-3-Clause. From 1.39.4, a patch release on 11 September 2026 whose notes include “Add Weaviate license and license switch”, code outside the wl directory is BSD-3-Clause and code inside it needs a separate enterprise licence enabled by a key. In 1.40.0 that directory holds backupdedupe, namespaces and selfrecovery, and the notes include “Make namespaces a feature requiring license”. The 1.38 line still carried the old file at 1.38.20. The core stays BSD; check whether the features you need sit inside wl.
The others, briefly
- OpenSearch and Elasticsearch. If you already run one, try its vector features first. Elasticsearch now defaults float vectors of 384 or more dimensions to
bbq_hnsw(binary quantisation) and smaller ones toint8_hnsw. OpenSearch’son_diskmode defaults to 32x compression with rescoring on. The fork story is in the earlier post on Elasticsearch and OpenSearch. - Redis and Valkey. Redis 8 has vector sets (
VADD,VSIM) under its RSALv2, SSPLv1 or AGPLv3 choice. valkey-search (BSD 3-Clause) offers HNSW and flat indexes with filters; 1.3.0-rc1 adds BM25 andFT.HYBRIDfusion and needs Valkey 9.1.0 or later. Good for small, hot sets, costly for 17 GiB of floats. - Chroma and LanceDB. Both Apache 2.0 and easy for prototypes and offline evaluation. LanceDB is embedded on Lance files; see the Chroma CVE below before exposing a Chroma server.
Filtering: where vector search quietly goes wrong
Real queries are “nearest 10 chunks for this brand, in Hindi, not archived”. Post-filtering runs the ANN search and then drops rows; filter-aware search applies the filter during graph traversal.
pgvector post-filters, and its README says so: if a condition matches 10% of rows, a default HNSW scan returns about 4 matching rows on average. Iterative scans (0.8.0 and later) keep scanning up to hnsw.max_scan_tuples, 20,000 by default. My drill with a 1% filter gave 0.4 rows per query on defaults and 10 rows with iterative scans. It also showed that adding a B-tree on the filter column did not change the plan: the planner still chose HNSW and returned too few rows. For selective filters, enable iterative scans, use partial HNSW indexes for a few large tenants, or force exact search on the small subset.
Qdrant filters during graph search. In my drill the same filter gave full recall either way, but a payload index on category cut p50 latency from 13.3 ms to 1.3 ms. Create payload indexes for every filtered field.
Hybrid search: BM25 plus vectors
Embeddings are weak on exact tokens such as order IDs, SKUs and brand names; keyword search is weak on meaning. Support content needs both. Reciprocal Rank Fusion (RRF) is the simple default: each result scores 1 / (k + rank) per list, summed, with k commonly 60.
The dedicated engines fuse natively. PostgreSQL’s built-in ts_rank and ts_rank_cd are not BM25; extensions such as ParadeDB’s pg_search (AGPL 3.0) and Tiger Data’s pg_textsearch (PostgreSQL License, PostgreSQL 17 and 18) add it. With built-in full-text search, fusion looks like this; I ran it on PostgreSQL 18.6 with pgvector 0.8.7:
WITH semantic AS (
SELECT id, row_number() OVER (ORDER BY dist) AS r
FROM (SELECT id, embedding::halfvec(768) <=> $1::halfvec(768) AS dist
FROM chunks WHERE tenant_id = $2
ORDER BY dist LIMIT 50) s
),
keyword AS (
SELECT id, row_number() OVER (ORDER BY rank DESC) AS r
FROM (SELECT id, ts_rank_cd(tsv, q) AS rank
FROM chunks, websearch_to_tsquery('english', $3) AS q
WHERE tenant_id = $2 AND tsv @@ q
ORDER BY rank DESC LIMIT 50) k
)
SELECT id, sum(1.0 / (60 + r)) AS rrf_score
FROM (SELECT * FROM semantic UNION ALL SELECT * FROM keyword) both_lists
GROUP BY id
ORDER BY rrf_score DESC
LIMIT 10;
Rank after the inner LIMIT so the vector side can still use the HNSW index. For Hinglish text the english configuration stems badly; test simple too.
Licences: read the files
The licence column in the first table comes from each LICENSE file, not from marketing pages. Most of this space is permissive: PostgreSQL License for pgvector and pgvectorscale, Apache 2.0 for Qdrant, Milvus, Chroma, LanceDB and OpenSearch k-NN, BSD for valkey-search. Four need a closer look: VectorChord (AGPLv3 or Elastic License v2), Weaviate since 1.39.4 (the wl boundary), ParadeDB pg_search (AGPL 3.0) and Redis 8 (RSALv2, SSPLv1 or AGPLv3). AGPL and source-available terms matter most if you may one day offer search as part of a hosted product.
Operations: the part demos skip
Day-two work differs more than features do. pgvector rides on PostgreSQL backups, PITR, streaming replication and ALTER EXTENSION vector UPDATE; it has no native vector sharding beyond partitioning or Citus. Qdrant has snapshots (local or S3), a replication factor per collection and a fixed shard count. Milvus keeps data in object storage and upgrades component by component. Weaviate has backup modules for S3, GCS, Azure or disk, and since 1.39.4 even a patch upgrade means re-reading the LICENSE file.
The hardest event is changing the embedding model. Vectors from two models are not comparable, so a new model means re-embedding every chunk and a new index. Plan for it:
- Store the model name and version with every vector.
- Write new vectors to a new column, table or collection while the old one serves.
- Compare both on a fixed set of real queries with known good answers.
- Switch reads (a view in PostgreSQL, an alias in Qdrant or Milvus), keep the old index for rollback, then drop it.
For 6 million chunks, embedding time and cost usually exceed index build time.
Security advisories
All IDs below were confirmed as PUBLISHED through the CVE Services API on 8 October 2026.
Project
CVE
What happened
Fixed in
pgvector
CVE-2026-3172
Parallel HNSW build overflow: data leak or crash (0.6.0 to 0.8.1)
0.8.2
pgvector
CVE-2026-18022
Integer wraparound in IVFFlat build; 32-bit systems only
0.8.6
pgvector
CVE-2026-103484
IVFFlat build out-of-bounds write, possible code execution
0.8.7
Qdrant
CVE-2024-3829
File read and write through symlinks in snapshot recovery
1.9.0
Qdrant
CVE-2026-25628
Append to arbitrary files via /logger with read-only access
1.16.0
Milvus
CVE-2025-64513
Unauthenticated bypass of Proxy authentication
2.4.24, 2.5.21, 2.6.5
Milvus
CVE-2026-26190
Port 9091 exposed the REST API without authentication
2.5.27, 2.6.10
Milvus
CVE-2026-69111
Unauthenticated /management/stop on port 9091
Record lists up to 2.6.22, and 3.0.0, as affected
Weaviate
CVE-2025-67818
Path traversal when restoring a backup
1.33.4
Weaviate
CVE-2026-59093
RBAC role assignment did not check the assigner’s permissions
1.38.0
Chroma
CVE-2026-45829
Pre-auth code injection via trust_remote_code
Record lists 1.0.0 and later; no fixed version given
The patterns: restore only snapshots you produced, keep management ports such as Milvus 9091 private, and remember that for pgvector “a database user” includes any application role that can run CREATE INDEX.
Benchmarks: whose run is it?
ANN-Benchmarks is the closest to neutral. Its README says current results are from April 2025 on an AWS r6i.16xlarge, each benchmark on a single CPU: good for comparing algorithms, weak for predicting a networked, filtered service.
VectorDBBench has filtered, streaming and full-text cases and an MIT licence, so you can run it yourself. Its README says it is sponsored by Zilliz, the company behind Milvus, and its leaderboard lives on zilliz.com, so treat that leaderboard as a vendor’s own run. Qdrant’s benchmark page is Qdrant’s own run, marked as updated in February 2023 and January/June 2024. The pgvectorscale and VectorChord figures are vendor claims too. That does not make them wrong; it makes them a hypothesis.
A fair in-house test uses your own embeddings and filters, a recall target fixed first (say 0.95 recall@10), real concurrency, and p95 rather than a mean.
A decision guide
Your situation
Reasonable first choice
Under about 10 million vectors, data already in PostgreSQL
pgvector with halfvec HNSW and iterative scans
Same, but the index no longer fits memory
Test pgvectorscale, or VectorChord after a licence review
Filter-heavy search, separate scaling, tens of millions of vectors
Qdrant
Hundreds of millions of vectors, platform team, object storage in place
Milvus
RAG needing strong built-in hybrid search, wl boundary acceptable
Weaviate
Already running OpenSearch or Elasticsearch
Their vector features first
For the Pune team, I would start with pgvector: a halfvec HNSW expression index (about 11 GiB by my estimate), iterative scans for brand and language filters, RRF hybrid search in SQL, and retrieval on a read replica, away from order traffic. Then write down two exit criteria: if replica p95 misses the target at real concurrency, or the corpus heads past about 30 million chunks, evaluate Qdrant on the same query set.
A practical checklist
- Fix recall and latency targets on real queries with known good answers.
- Do the memory arithmetic for vectors, index, replicas and two years of growth.
- Choose the embedding dimension deliberately; fewer dimensions shrink every number here.
- Store the model name and version with every vector.
- Use
halfvecor quantisation, and measure recall with rescoring on and off. - Test filtered queries at real selectivity and check they return the full k rows.
- Add hybrid search for IDs, codes and brand names.
- Put the vector store on the patch calendar; pgvector alone had three CVEs this year.
- Lock down network exposure and auth before loading real data.
- Rehearse a re-embedding with a parallel index and a switch.
Common mistakes
- Trusting defaults on filtered queries. pgvector can return fewer rows than
LIMIT; enable iterative scans and check counts. - Assuming quantisation always rescores. Qdrant’s scalar int8 does not by default; in my drill that cost about 4 points of recall.
- Forgetting the vector copy in the index. Size pgvector’s table and HNSW index separately.
- Mixing embeddings from two models in one index. Distances become meaningless and nothing errors.
- Leaving management ports open. Several Milvus CVEs needed only network access to port 9091.
Drill: pgvector and Qdrant on one VM
No root and no containers. On 8 October 2026, from about 6:24 AM to 6:27 AM IST, on a shared Linux VM with 8 vCPUs (Intel Xeon) and 15.6 GiB of RAM, I ran PostgreSQL 18.6 with pgvector 0.8.7 and Qdrant 1.19.2.
- pgvector:
postgresql-18-pgvector_0.8.7-1.pgdg13+1_amd64.deb(274,800 bytes) from the PostgreSQL apt repository, SHA-256 matching the signed apt index, unpacked into my own directory and loaded through PostgreSQL 18’s newextension_control_pathanddynamic_library_pathsettings. - Qdrant:
qdrant-x86_64-unknown-linux-gnu.tar.gz(32,844,429 bytes) from GitHub. There is no checksum file; the SHA-256 matched the digest GitHub shows for the asset. It ran on 127.0.0.1 with telemetry off. - Data: 100,000 synthetic 384-dimension vectors around 2,000 random cluster centres, normalised (146.5 MiB raw), and 200 queries from the same distribution. Ground truth was exact cosine top 10 from NumPy. A
categoryfrom 0 to 99 per row;category = 7matched 989 rows, about 1%.
pgvector
COPY loaded the table in 10.6 s. With maintenance_work_mem = 1GB and 4 parallel workers, the HNSW index (m = 16, ef_construction = 64) built in 10.3 s.
Query
recall@10
Rows (avg)
p50
p95
Exact, no index
1.000
10
26.98 ms
30.48 ms
HNSW, ef_search = 10
0.924
10
0.46 ms
0.76 ms
HNSW, ef_search = 40 (default)
0.990
10
0.73 ms
1.27 ms
HNSW, ef_search = 100
1.000
10
1.75 ms
2.37 ms
Filter, exact (before the index)
1.000
10
12.70 ms
14.33 ms
Filter, HNSW, iterative off
0.041
0.4
0.88 ms
1.26 ms
Filter, HNSW, relaxed_order
0.841
10
18.46 ms
30.17 ms
Filter, HNSW, strict_order
0.665
10
22.58 ms
36.83 ms
Filter, plus B-tree on category, iterative off
0.041
0.4
0.89 ms
1.38 ms
The table was 166,436,864 bytes (159 MB), the float32 HNSW index 204,808,192 bytes (195 MB), and a halfvec HNSW expression index 117,039,104 bytes (112 MB), built in 9.3 s. With the B-tree present, EXPLAIN still chose the HNSW index with Filter: (category = 7).
Qdrant
Upload in batches of 1,000 took 26.4 s; at 27.4 s the collection was green, with 97,000 vectors in HNSW segments and the rest in a small segment searched exactly. Defaults: m 16, ef_construct 100.
Query
recall@10
p50
p95
exact: true
1.000
4.80 ms
7.46 ms
hnsw_ef = 10
0.987
1.43 ms
1.81 ms
hnsw_ef = 40
1.000
1.60 ms
1.98 ms
Filter, no payload index
1.000
13.29 ms
18.24 ms
Filter, integer payload index
1.000
1.27 ms
1.76 ms
Scalar int8, default params (hnsw_ef 40 or 100)
0.958
1.46 to 1.57 ms
1.83 to 1.94 ms
Scalar int8, rescore: true
1.000
1.51 ms
2.00 ms
Memory from /proc: 58 MiB RSS empty, 1,279 MiB after indexing, peak (VmHWM) 1,427 MiB, nearly ten times the raw vectors. After int8 quantisation and the rebuild, RSS settled near 318 MiB (94 MiB anonymous, 223 MiB file-backed). Storage was 271 MB on disk; a snapshot was 283,015,680 bytes. My loopback-only run did not print the new no-API-key warning; I did not test other bindings.
Honest limits: the data is synthetic and clustered, so recall will not transfer to real embeddings. Latencies are single-client and include client overhead (psycopg over a Unix socket, JSON over HTTP for Qdrant), so compare settings within an engine, not engines with each other. 200 queries on a shared VM is a small sample; I did not test concurrency, Milvus, Weaviate, clusters, upgrades or 6 million vectors. Nothing was blocked by an approval check.
What to unlearn and re-learn
- Unlearn “AI search needs an AI database”. Re-learn that nearest-neighbour search is an index type, and PostgreSQL has a good one.
- Unlearn “approximate means slightly wrong”. Re-learn that with filters, approximate can mean empty; check row counts.
- Unlearn “more RAM fixes vector search”. Re-learn the arithmetic of dimensions, precision, replicas and rescoring.
- Unlearn “licence checked once”. Re-learn to read the LICENSE file on every upgrade; Weaviate changed it in a patch release.
- Unlearn “embeddings are permanent”. Re-learn that a model change is a full re-index, and design for two indexes at once.
Revisit your vector search before you add a vector database
The Pune team does not need to pick a winner from a slide deck. Learn how HNSW, IVF, DiskANN and quantisation spend memory and recall. Unlearn the idea that a new kind of query needs a new kind of database. Re-learn filtering, hybrid ranking and re-embedding as ordinary engineering. Practise on a copy of real data with a recall target, as small as my 100,000-vector drill or as large as the full 6 million. Then apply the simplest setup that meets the target, write down the numbers that would justify moving, and revisit the choice when those numbers change.
Sources
- pgvector repository, 0.8.7 release, README, CHANGELOG, LICENSE, issue 1036, issue 959, pgvector-python RRF example
- CVE-2026-3172, CVE-2026-18022, CVE-2026-103484
- pgvectorscale repository, 0.9.1 release, 0.2.0 release, LICENSE, Tiger Data: pgvector is now faster than Pinecone (vendor benchmark)
- VectorChord repository, 1.1.1 release, 1.1.0 release, LICENSE
- PostgreSQL 18: client connection defaults, extension_control_path, PostgreSQL text search controls, ParadeDB repository, ParadeDB LICENSE, ParadeDB 0.26.0 release, pg_textsearch repository, pg_textsearch LICENSE
- Qdrant repository, 1.19.2 release, 1.19.0 release, 1.0.0 release, LICENSE
- Qdrant docs: quantization, indexing, filtering, hybrid queries, collections and aliases, security, snapshots, distributed deployment, Qdrant benchmarks (vendor run)
- CVE-2024-3829, CVE-2026-25628
- Milvus repository, 3.0.2 release, 3.0.0 release, 2.6.25 release, 1.0.0 release, 2.0.0 release, LICENSE
- Milvus docs: release notes, architecture overview, index explained, authentication, full text search, aliases, LF AI & Data: Milvus graduation
- CVE-2025-64513, CVE-2026-26190, CVE-2026-69111
- Weaviate repository, 1.40.0 release, 1.39.10 release, 1.39.4 release, 1.0.0 release, LICENSE at 1.40.0, LICENSE at 1.39.3, LICENSE at 1.38.20, wl directory
- Weaviate docs: vector index, vector index concepts, vector quantization, hybrid search, CVE-2025-67818, CVE-2026-59093
- Chroma repository, 1.5.9 release, 1.0.0 release, LICENSE, chromadb on PyPI, CVE-2026-45829
- LanceDB repository, 0.40.0 release, LICENSE
- valkey-search repository, 1.3.0-rc1 release, 1.2.1 release, 1.0.0 release, LICENSE
- Redis repository, 8.10.2 release, 8.0.0 release, LICENSE.txt, Redis vector sets docs
- Elasticsearch dense_vector reference, Elasticsearch 9.5.5 release, Elasticsearch LICENSE.txt, OpenSearch disk-based vector search, OpenSearch k-NN 3.9.0.0 release, k-NN LICENSE
- HNSW paper, arXiv 1603.09320, Billion-scale similarity search with GPUs, arXiv 1702.08734, Engineering at Meta: Faiss announcement, FAISS repository, FAISS 1.15.1 release, FAISS LICENSE, DiskANN, NeurIPS 2019, DiskANN repository
- ANN-Benchmarks site, ANN-Benchmarks repository and README, ANN-Benchmarks paper, arXiv 1807.05614, VectorDBBench repository (Zilliz), VectorDBBench 2.0.0 release, Zilliz leaderboard (vendor run)
- CVE Services API, PostgreSQL apt repository
Releted Posts
PgBouncer and its alternatives: revisit your PostgreSQL connection pooler
A team runs about forty services on Kubernetes against one PostgreSQL primary. During a sale, the autoscaler adds pods, each opens its own pool of ten connections, and PostgreSQL starts refusing logins with “remaining connection slots are reserved”.
Read morePrometheus long-term storage: revisit Thanos, Mimir and VictoriaMetrics before you scale
A platform team runs two Prometheus servers per cluster as an HA pair, with 15 days of local retention. In one quarter, three requests arrive: the SRE lead wants a year of history for capacity planning, a product team adds a customer_id label to a request metric, and the auditors ask what happens to metrics when a disk dies.
Read morePostgreSQL asynchronous I/O: revisit your settings before you upgrade
A team upgrades its reporting database from PostgreSQL 16 to 18 over a weekend. On Monday the nightly scans finish a little faster, and nobody knows why, because nobody changed a setting.
Read more