PostgreSQL asynchronous I/O: revisit your settings before you upgrade
A team upgrades its reporting database from PostgreSQL 16 to 18 over a weekend. On Monday the nightly scans finish a little faster, and nobody knows why, because nobody changed a setting. A month later someone sets io_method to io_uring in a new container image because a blog said it is faster, and it does not work inside the container. Both stories have one root: PostgreSQL 18 changed how it reads data, and old I/O settings now mean something different.
This post covers why PostgreSQL relied on synchronous reads for so long, what the asynchronous I/O (AIO) subsystem in PostgreSQL 18 changes and what it does not, the three io_method choices, the changed settings, how to see AIO in pg_aios, EXPLAIN and pg_stat_io, what PostgreSQL 19 changes, an upgrade checklist, common mistakes and a sandbox drill with numbers from my own run. Facts come from the PostgreSQL release notes, documentation, source README and commit messages, plus the Docker, containerd and Linux kernel docs, all checked on 6 October 2026.
The short version
- PostgreSQL 18.0 was released on 25 September 2025 and is the current major version (18.6 is the current minor). PostgreSQL 19 is in beta: Beta 4 came out on 24 September 2026, and the project says GA “may also occur in October”.
- Before 18, every data file read was a synchronous system call, helped by kernel readahead and some posix_fadvise() hints.
- PostgreSQL 18 adds an AIO subsystem that lets a backend queue several read requests at once. The release announcement names sequential scans, bitmap heap scans and vacuum as supported operations.
io_methodpicks how AIO runs:worker(the default, all platforms),io_uring(Linux only, needs a build with liburing) orsync(the old style behaviour). It can only be changed with a restart.- The defaults of
effective_io_concurrencyandmaintenance_io_concurrencywent up to 16 in PostgreSQL 18. In 17 they were 1 and 10. Old explicit values in your config will override the new defaults. - AIO in 18 is for reads. Checkpointer, background writer and backend writes, and WAL writes, still use the existing synchronous path.
- The project’s announcement says benchmarking “has demonstrated performance gains of up to 3x in certain scenarios”. That is the project’s own measurement. On my sandbox the gains were much smaller, as you will see.
- In PostgreSQL 19 (beta),
io_workersis replaced by a self-sizing pool (io_min_workers,io_max_workersand two timing settings), and EXPLAIN gets anIOoption.
Why PostgreSQL read data synchronously for so long
PostgreSQL uses buffered I/O. When a backend needs a page that is not in shared_buffers, it calls read(), waits for the kernel to return the block, and only then continues. The kernel’s page cache sits underneath, and its readahead notices sequential patterns and fetches the next part of the file early. For sequential scans on local disks this worked well enough for years.
The weak spots were access patterns the kernel cannot guess. A bitmap heap scan jumps between pages that are in order but not adjacent. For such cases PostgreSQL issued posix_fadvise(POSIX_FADV_WILLNEED) hints for blocks it knew it would need soon, controlled by effective_io_concurrency. The PostgreSQL 17 docs say this setting “only affects bitmap heap scans” and that it “depends on an effective posix_fadvise function, which some operating systems lack”.
The commit that added the core AIO infrastructure lists the limits plainly: posix_fadvise() prefetching has no completion feedback, doubles the system calls, only works with buffered I/O and only on some operating systems. It also names a second goal, direct I/O, which is not usable for real workloads while every I/O is synchronous. The 18 announcement puts it in user terms: operating systems “lack insight into database-specific access patterns”.
What asynchronous I/O changes in PostgreSQL 18
The release notes describe the feature as allowing “backends to queue multiple read requests, which allows for more efficient sequential scans, bitmap heap scans, vacuums, etc.” Three ideas matter for a working engineer:
- Reads are issued ahead and in parallel. Code that reads through the read stream interface (sequential scans, bitmap heap scans, vacuum, ANALYZE and others) can have several reads in flight while it works on pages that already arrived. The core commit says moving read streams to AIO converts “all read stream users” at once. Data still goes through the OS page cache.
- Reads are combined. Adjacent blocks are merged into one larger read, up to
io_combine_limit(default 128 kB, which is 16 pages of 8 kB). In my drill, a scan of 242,425 blocks showed up as 15,156 reads in pg_stat_io, which is about 128 kB per read. - The issuing process does not do all the waiting. With
worker, background I/O worker processes run the system calls. Withio_uring, the backend submits requests to the kernel through a shared ring and collects completions later.
What is still synchronous in 18
The AIO core commit says patches to use AIO for the checkpointer and background writer were “reasonably close to being ready”, and that there were prototypes for WAL, relation extension and backend writes. None of these are in 18. So:
- Dirty page writes by the checkpointer, background writer and backends are synchronous as before.
- WAL writes and flushes are synchronous as before. Commit latency on a write heavy system will not change because of
io_method. - Plain index scans do not get read-ahead for their heap fetches. The project’s AIO wiki page (last edited 28 July 2025) lists “readahead of the table fetches for ordered index scans” as work still to be done.
If your workload is mostly index lookups on a hot working set, expect little difference. Large scans, bitmap scans and vacuum on uncached data are where AIO shows up.
The three io_method choices
io_method can only be set at server start. The default is worker.
# postgresql.conf (PostgreSQL 18)
io_method = worker # worker | io_uring | sync, restart needed
io_workers = 3 # only used with worker, reload is enough
worker (default)
The source README says worker “is available on every platform postgres runs on”. The backend pushes I/O requests into a shared memory queue, and I/O worker processes run ordinary synchronous system calls and complete the I/O. From the backend’s point of view the I/O is asynchronous. You will see these processes in ps as postgres: io worker 0, io worker 1 and so on.
io_workers sets the pool size, default 3, and can be changed with a reload. The commit that added worker mode notes that some I/Os cannot be run by another process and are done synchronously by the backend.
Fits: most installations, containers with a default seccomp profile, non Linux systems, and places where you do not control the kernel or the build.
io_uring
io_uring is a Linux asynchronous I/O interface built on ring buffers shared between user space and the kernel. The README says PostgreSQL’s io_uring method needs Linux 5.1 or later and, unlike worker mode, “dispatches all IO from within the process, lowering context switch rate / latency”. Its commit says it can be “considerably faster” than worker when many small I/Os are issued, as worker context switches add up and the worker count can become the limit.
It needs a PostgreSQL build with liburing: --with-liburing for configure or -Dliburing for meson. To check your binary:
pg_config --configure | tr ' ' '\n' | grep -i uring
ldd "$(pg_config --bindir)/postgres" | grep -i uring
The PGDG Debian package of 18.6 that I installed is built with liburing. Do not assume every image or managed service is.
io_uring is also often blocked on purpose:
- Docker’s default seccomp profile blocks
io_uring_setup,io_uring_enterandio_uring_register. Docker’s docs say these are “blocked due to security vulnerabilities that can be exploited to break out of containers” (moby/moby pull request 46762). - containerd merged a similar change to its RuntimeDefault seccomp profile (containerd pull request 9320), which matters for Kubernetes pods that use RuntimeDefault.
- The Linux sysctl
kernel.io_uring_disabledcan disable io_uring creation for unprivileged processes not inio_uring_group(value 1) or for all processes (value 2). The default is 0.
Fits: Linux hosts or VMs you control, a liburing build, no seccomp or sysctl block, and workloads with many concurrent reads where you have measured a benefit.
sync
The docs describe sync as executing “asynchronous-eligible I/O synchronously”, and the README calls it useful for debugging. Use it as a fallback if you suspect AIO in a problem, or as the baseline in your own tests.
Storage changes the answer
The PostgreSQL 18 docs say higher effective_io_concurrency values “will have the most impact on higher latency storage where queries otherwise experience noticeable I/O stalls and on devices with high IOPs”. In practice:
- Cloud network block storage has higher per request latency. More reads in flight hide it, so AIO and the concurrency settings matter most here.
- Local NVMe has low latency. Gains per query are smaller, and CPU may become the limit.
- Data already in the OS page cache is a memory copy, not a disk read. My drill below shows that methods behave quite differently in this case.
The settings that changed
All values below are from the PostgreSQL 18 documentation unless marked otherwise.
- effective_io_concurrency: concurrent storage I/O operations a session tries to start. Range 1 to 1000, or 0 to disable. Default 16 (1 in 17). Can be set per tablespace.
- maintenance_io_concurrency: the same for maintenance work such as vacuum. Default 16 (10 in 17). Can be set per tablespace.
- io_combine_limit: largest I/O size for combined reads. Default 128 kB. If set above
io_max_combine_limit, the lower value is used silently. - io_max_combine_limit: new in 18, server start only, default 128 kB, caps
io_combine_limit. The docs say the maximum possible size is typically 1 MB on Unix and 128 kB on Windows. To use larger reads, raise both. - io_max_concurrency: maximum I/Os one process can have in flight. The default -1 picks a value from
shared_buffersand the maximum number of processes, capped at 64. Server start only. On my sandbox it resolved to 64. - io_method and io_workers: as above.
The commit that raised the default explains why: 1 was a conservative choice from when the setting was introduced, and tests on high latency cloud storage and local NVMe showed even slightly higher values improved timings substantially. It notes that 1 performed worse than 0, because it added system calls without prefetching enough.
-- See what your server is really using, and where each value came from
SELECT name, setting, unit, source, sourcefile
FROM pg_settings
WHERE name IN ('io_method', 'io_workers', 'effective_io_concurrency',
'maintenance_io_concurrency', 'io_combine_limit',
'io_max_combine_limit', 'io_max_concurrency')
ORDER BY name;
-- Per tablespace override, for example on slower storage
ALTER TABLESPACE archive_ts SET (effective_io_concurrency = 64);
If source says configuration file for effective_io_concurrency, an old value is overriding the new default. Decide on purpose whether to keep it.
How to see AIO in action
pg_aios
The new pg_aios view has one row per AIO handle in use: process ID, state (such as SUBMITTED), operation (readv or writev), offset, length, target and a synchronous flag. The docs call it “mainly useful for developers”, and by default only superusers and pg_read_all_stats members can read it. Rows exist only while I/O is in flight, so query it during a big scan:
SELECT pid, state, operation, off, length, target_desc, f_sync
FROM pg_aios
LIMIT 10;
In my drill, during a cold scan with worker, it showed several SUBMITTED readv rows of 131,072 bytes each (for example “blocks 55455..55470”), with f_sync false: 16 combined blocks per read, several in flight.
EXPLAIN (ANALYZE, BUFFERS) and track_io_timing
In PostgreSQL 18, EXPLAIN ANALYZE includes buffer information automatically, and with track_io_timing = on it shows I/O timings. With AIO that timing is how long the backend waited for reads, not total disk time, because reads overlap with work. Compare both “Execution Time” and “I/O Timings” across methods.
SET track_io_timing = on; -- needs superuser or the right privilege
EXPLAIN (ANALYZE, BUFFERS) SELECT count(*) FROM big_table;
pg_stat_io
pg_stat_io shows cluster wide I/O counts by backend type, object and context. In 18 it gained read_bytes, write_bytes and extend_bytes (replacing op_bytes) and WAL rows. Bytes divided by reads gives the average read size, which shows whether read combining works:
SELECT backend_type, object, context, reads,
pg_size_pretty(read_bytes) AS read_bytes,
round(read_bytes / NULLIF(reads, 0)) AS avg_read_bytes,
round(read_time) AS read_ms
FROM pg_stat_io
WHERE reads > 0
ORDER BY reads DESC;
Wait events help too. Under IO you will find AioIoCompletion (waiting for another process to complete I/O), AioIoUringSubmit and AioIoUringExecution. Under LWLock, AioWorkerSubmissionQueue; under Activity, IoWorkerMain for an I/O worker waiting in its main loop.
What changes in PostgreSQL 19
PostgreSQL 19 is in beta, so details can still change. The current beta release notes list three AIO related items:
- “Improve asynchronous I/O read-ahead scheduling for large requests.”
io_method = workernow sizes its pool automatically. The commit saysio_workers“is replaced with”io_min_workers(default 2),io_max_workers(default 8, up to 32),io_worker_idle_timeout(default 1 minute) andio_worker_launch_interval(default 100 ms). If you setio_workersin 18, plan to replace it when you move to 19.- A new
EXPLAIN (ANALYZE, IO)option reports prefetch queue distance, number of I/O requests, average request size, I/O waits and average concurrent requests for scan nodes.
The 19 notes also add streaming reads to hash index bulk deletion, GIN index vacuuming, bloom indexes and pgstattuple, so more paths can use AIO. I found no item for asynchronous writes or WAL. The io_method choices and the defaults of 16 are unchanged in the 19 docs.
What to unlearn and re-learn
- Unlearn “PostgreSQL leaves read-ahead to the kernel”. From 18 it plans its own reads on read stream paths. Re-learn that the page cache is still underneath, and cached data behaves differently by method.
- Unlearn “effective_io_concurrency only matters for bitmap heap scans and for RAID”. Re-learn it as the per session read depth for AIO, and set it per tablespace if storage differs.
- Unlearn “asynchronous means writes are faster too”. In 18 it is reads. Commit latency is still about WAL and fsync.
- Re-learn to measure with
pg_stat_iobytes per read and wait events, not only query time.
Upgrade checklist
- Read your old
postgresql.confand anyALTER SYSTEMvalues foreffective_io_concurrencyandmaintenance_io_concurrency. Remove them to get the new default of 16, or keep them with a written reason. - Check the build:
pg_config --configurefor liburing if you wantio_uring. - Check the runtime: seccomp profile (Docker default and containerd RuntimeDefault block io_uring),
kernel.io_uring_disabled, and your managed service’s allowed values. - Start on the default
worker. Treatio_uringas an experiment you must win with numbers on your own storage. - Keep
track_io_timingon for the test period (check its overhead withpg_test_timingfirst), and capturepg_stat_iobefore and after. - Re-check vacuum duration and I/O load, because
maintenance_io_concurrencywent up from 10 to 16 by default. - If you raise
io_combine_limit, raiseio_max_combine_limittoo, and restart. - For 19, plan to replace
io_workerswithio_min_workersandio_max_workers.
Common mistakes
- Copying io_method = io_uring into a container image. Docker’s default seccomp profile blocks it. Test where you deploy.
- Keeping effective_io_concurrency = 1 from an old template. It silently cancels the new default, and the project found 1 worse than 0.
- Setting very high concurrency on shared or throttled storage. The docs warn that “unnecessarily high values may increase I/O latency for all queries on the system”.
- Raising io_combine_limit alone. It is capped silently by
io_max_combine_limit. - Benchmarking with a warm page cache and calling it a disk test. Evicting shared buffers alone is not enough. The OS cache still serves the data.
- Expecting AIO to fix write or commit latency. It does not in 18.
- Quoting “up to 3x” as a promise. It is the project’s measurement for certain scenarios. Your hardware decides.
Drill: compare sync, worker and io_uring on one cold scan
You can do this in an hour on a sandbox. Do not do it on a shared production host.
Step 1. Install PostgreSQL 18 with liburing support, create a throwaway cluster and set:
shared_buffers = 256MB
max_parallel_workers_per_gather = 0 # isolate one backend
track_io_timing = on
Step 2. Create a table well above shared_buffers, freeze it and checkpoint:
CREATE EXTENSION pg_buffercache;
CREATE TABLE t (id bigint, pad text) WITH (autovacuum_enabled = off);
INSERT INTO t SELECT g, repeat('x', 200) FROM generate_series(1, 8000000) g;
VACUUM (FREEZE, ANALYZE) t;
CHECKPOINT;
SELECT pg_relation_filepath('t'), pg_size_pretty(pg_relation_size('t'));
Step 3. For each method: set io_method with ALTER SYSTEM, restart, evict the table from shared buffers with SELECT pg_buffercache_evict_relation('t'); (new in 18, superuser only), then from the OS page cache. As root on a dedicated test machine, echo 1 > /proc/sys/vm/drop_caches does it. Otherwise call posix_fadvise(POSIX_FADV_DONTNEED) on the table’s files and check residency with fincore or mincore().
Step 4. Run EXPLAIN (ANALYZE, BUFFERS, TIMING OFF) SELECT count(*) FROM t; and note both timings. Run again without the OS eviction to see the page cache case. Repeat at least five times, interleaved, and use medians.
Step 5. While one cold scan runs, query pg_aios from another session, and afterwards check pg_stat_io bytes per read.
My run, and its limits
I ran this on 6 October 2026 on my sandbox: a shared VM container with 8 vCPUs, 15 GB RAM, Debian 13, Linux 6.12, a virtual disk under an overlay filesystem, and PostgreSQL 18.6 from the PGDG apt repository. The table was 1,894 MB (242,425 blocks). I could not drop caches globally (permission denied on /proc/sys/vm/drop_caches), so I used posix_fadvise(DONTNEED) on the table’s files and confirmed 0% residency with mincore() before each cold run. The host may still have cached the virtual disk: cold reads ran at well over 1 GB/s, which is not what cloud block storage looks like. Other workloads shared the machine. These are one box’s numbers, not a benchmark.
Median of five interleaved rounds, in milliseconds (I/O wait from “I/O Timings: shared read”):
io_method
Cold: execution
Cold: I/O wait
OS cache warm: execution
OS cache warm: I/O wait
sync
1326
852
995
585
worker
1264
766
529
18
io_uring
1126
508
840
382
What I saw:
- Cold,
io_uringwas about 15% faster thansyncandworkerabout 5% faster, with noisy rounds (worker ranged from 1159 to 1782 ms). Nothing near 3x, which is expected on fast, low latency storage. - With the data in the OS page cache but not in shared buffers,
workerwas close to 2x faster thansyncand its I/O wait almost vanished. My reading, not profiled, is that the three I/O workers did the page cache copies in parallel with the backend, while withio_uringthe backend did more of that work itself. - On
io_uring, cold, three rounds each:effective_io_concurrency1 gave a median of 1380 ms, 16 gave 1152 ms and 64 gave 1091 ms. pg_stat_ioshowed 15,156 reads for 1,894 MB, about 128 kB per read, matchingio_combine_limit.
Your result on cloud storage or NVMe will differ. That is the point of running it.
Revisit your I/O settings
PostgreSQL 18’s asynchronous I/O is a quiet change. It needs no extension or new SQL, and the default worker method works almost everywhere. That makes it easy to miss what changed: old concurrency values that now cap read depth, io_uring choices that a container profile blocks, and benchmarks that measure the page cache instead of the disk. Learn what AIO covers today (reads), unlearn the habit of leaving read-ahead entirely to the kernel, re-learn your tuning with pg_stat_io and pg_aios, practise the drill on your own storage, and apply only the settings your numbers support. Then revisit it again when PostgreSQL 19 ships with its self-sizing I/O worker pool.
Sources
- PostgreSQL 18 Released! (project announcement, AIO section)
- PostgreSQL 18.0 release notes
- PostgreSQL 18 docs: Resource Consumption, I/O settings
- PostgreSQL 17 docs: Resource Consumption (old effective_io_concurrency text and defaults)
- PostgreSQL 18 docs: pg_aios view
- PostgreSQL 18 docs: Cumulative statistics, pg_stat_io and wait events
- PostgreSQL 18 docs: EXPLAIN
- PostgreSQL 18 docs: pg_buffercache
- PostgreSQL 18 docs: Building with Meson (-Dliburing)
- PostgreSQL 18 docs: Building with Autoconf and Make (–with-liburing)
- PostgreSQL source: src/backend/storage/aio/README.md (REL_18_STABLE)
- Commit da7226993: aio: Add core asynchronous I/O infrastructure
- Commit 247ce06b8: aio: Add io_method=worker
- Commit c325a7633: aio: Add io_method=io_uring
- Commit ff79b5b2a: Increase default effective_io_concurrency to 16
- PostgreSQL wiki: AIO
- PostgreSQL 19 (beta) release notes
- PostgreSQL 19 (beta) docs: Resource Consumption, I/O worker settings
- PostgreSQL 19 (beta) docs: EXPLAIN IO option
- Commit d1c01b79d: aio: Adjust I/O worker pool automatically
- PostgreSQL 19 Beta 4 Released!
- PostgreSQL versioning policy (supported versions and current minors)
- Docker docs: Seccomp security profiles
- moby/moby pull request 46762: block io_uring syscalls in default seccomp profile
- containerd pull request 9320: disallow io_uring syscalls in RuntimeDefault
- Linux kernel docs: sysctl kernel, io_uring_disabled and io_uring_group
- Linux kernel docs: sysctl vm, drop_caches
- io_uring(7) Linux manual page
Releted Posts
Redis, Valkey, and Dragonfly benchmarks: read the method first
A benchmark number without its method sounds complete, but it does not tell you what happened. The first note in this series, Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache, covered licences, history, data file compatibility, and cluster mode.
Read moreMongoDB 8.0 Performance is 36% higher, but there is a catch…
TLDR: If your app is performance critical, think twice, thrice before upgrading to MongoDB 7.0 and 8.0. Here is why…
Read more
Optimizing MongoDB Performance with Indexing Strategies
MongoDB is a popular NoSQL database that efficiently handles large volumes of unstructured data. However, as datasets grow, query performance can degrade due to increased document scans.
Read more