Plate 60
fsync vs fdatasync vs none: Durable Write Cost Lab
Hands-on fsync vs fdatasync vs none: per-write and every-N sync on 256B and 4KiB records. Real p50/p95 show about 100-200x durability tax on localhost
Aditya Challa7 min read
Intro — what this post promises
Calling write() is not the same as durable data. The kernel is happy to park your bytes in the page cache and return. Crash (or power loss) later and those “writes” never happened.
This is a hands-on lab with measured numbers:
- Wall time for batches of small (256 B) and medium (4 KiB) records with no sync.
- Same batches with
fdatasyncper write andfsyncper write. - Amortization: sync every 10 / every 50 writes.
- Optional
O_SYNCopen — durability on everywrite(2). - Honest limits on an overlay/virtio lab box (not a battery-backed WAL appliance).
Related links:
- mmap vs read / O_DIRECT localhost lab
- sendfile vs userspace copy localhost lab
- pipe vs tmpfile IPC localhost lab
- nice / ionice CPU and disk priority lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 writing under /tmp on overlay atop virtio. No Docker. No public bind. Affiliates: 0. This measures sync cost, not crash recovery or fsck after echo b > /proc/sysrq-trigger.
Verdict up front: none → any per-write sync was the cliff (~100–200×). On this overlay, fsync ≈ fdatasync (within noise). Syncing every 50 medium records bought back hundreds of MB/s.
What fsync / fdatasync / O_SYNC actually change
| Knob | What it waits for (simplified) | Typical use |
|---|---|---|
none (buffered write) | bytes in page cache | temp files, best-effort logs |
fdatasync | file data (and needed metadata for that data) on storage | WAL / record append without full inode flush folklore |
fsync | data and metadata | rename-over, directory durability stories |
O_SYNC / O_DSYNC | each write behaves like write+sync | simple but brutal APIs |
Textbook claim: fdatasync is cheaper than fsync because it can skip some metadata. On this lab FS the gap was tiny — report that honestly instead of pasting a textbook 2×.
Related links:
Lab topology
Script: lab-evidence/32-fsync-vs-fdatasync/results/run_lab.py. Each arm creates/truncates a file, writes N fixed payloads, applies the sync policy, closes, records perf_counter wall time for the whole batch.
RPS / MB/s below are from p50 batch wall (ops and bytes ÷ p50 seconds) so a slow repeat does not silently dilute the headline.
Arm A — small records (256 B × 2000)
| Arm | p50 batch | p95 | per-rec p50 | RPS (p50) | MB/s (p50) |
|---|---|---|---|---|---|
| none | 1.31 ms | 2.72 | 0.65 µs | ~1.53M | ~373 |
| fdatasync / 1 | 259 ms | 268 | 130 µs | ~7.7k | ~1.9 |
| fsync / 1 | 240 ms | 261 | 120 µs | ~8.3k | ~2.0 |
| O_SYNC (500 rec) | 60.4 ms | 62.7 | 121 µs | ~8.3k | ~2.0 |
none → fdatasync/1 ≈ 198× on batch p50. That is the durability tax in one table.
fsync/1 was not slower than fdatasync/1 here (slightly faster on this run). Treat them as tied on overlay, not as a portable “always use fdatasync” win.
Stability×3 for none p50 (ms): 1.18 / 1.26 / 1.18. For fdatasync/1: 246 / 234 / 233.
Related links:
Arm B — medium records (4 KiB × 500)
| Arm | p50 batch | p95 | per-rec p50 | RPS (p50) | MB/s (p50) |
|---|---|---|---|---|---|
| none | 0.84 ms | 1.49 | 1.7 µs | ~593k | ~2317 |
| fdatasync / 1 | 75.1 ms | 88.7 | 150 µs | ~6.7k | ~26 |
| fsync / 1 | 93.5 ms | 99.3 | 187 µs | ~5.3k | ~21 |
| O_SYNC (200 rec) | 35.3 ms | 37.6 | 176 µs | ~5.7k | ~22 |
Medium none looks like memory copy (multi‑GB/s class) because it mostly dirties cache. Per-write sync drops you into the ~20–26 MB/s band on this box — still faster than the 256 B sync arms in MB/s (larger payload per barrier) but nowhere near “disk sequential” marketing numbers.
Stability×3 for med fsync/1 p50 (ms): 85.0 / 68.6 / 77.5 (noisy — quote bands, not fake precision).
Arm C — amortize: every 10 / every 50
Same 4 KiB × 500 shape; sync only every N records (plus a final sync if the last group is short).
| Arm | p50 | per-rec p50 | MB/s (p50) | vs sync/1 |
|---|---|---|---|---|
| fdatasync / 10 | 16.0 ms | 31.9 µs | ~122 | ~4.7× vs /1 |
| fsync / 10 | 12.8 ms | 25.6 µs | ~152 | ~7.3× |
| fdatasync / 50 | 5.70 ms | 11.4 µs | ~343 | ~13× |
| fsync / 50 | 5.24 ms | 10.5 µs | ~373 | ~18× |
Group commit is not a metaphor: every-50 recovered into the hundreds of MB/s while still issuing durability barriers. Databases and queues do this for a reason.
Related links:
How to read these numbers
none≠ durable. Fast path is page cache. Fine for scratch; wrong for “acked to disk.”- Per-write sync turns throughput into barrier rate. Small records hurt more in MB/s.
fsyncvsfdatasync: textbook gap did not show on this overlay run — do not ship a fake 3× claim from this post.O_SYNC: per-record cost matched sync/1 (~120–180 µs). Convenient API; same economic cliff.- Every-N is the practical dial between RPO fantasy and RPS reality.
Pitfalls we hit (or avoided)
- Reporting sum-of-repeats RPS — early harness bug understated RPS by ×repeats; numbers above use p50 batch wall.
- Assuming fdatasync always wins — not on this FS snapshot.
- Calling none “disk throughput” — it is not; see mmap / O_DIRECT labs for device-ish paths.
- One p50 without p95 — sync arms were tighter relative to mean; still show both.
- Claiming crash safety — we timed syscalls; we did not pull the plug.
Practical checklist
- Decide the durability contract (lost last N records OK? or not?).
- If you need sync, batch (every N / group commit /
fdatasyncon a WAL segment) before blaming the SSD. - Benchmark your FS — overlay ≠ XFS on NVMe ≠ cloud block with flaky flush.
- Pair with p95 and a clear record size; do not quote only MB/s from
none. - Keep evidence next to any “we fsync every message” capacity plan.
Related links:
- stdout buffering line vs full localhost lab
- context-switch pipe ping-pong localhost lab
- sendfile vs userspace copy localhost lab
Verdict
On this box, buffered none wrote 256 B records at 373 MB/s~~ (0.65 µs/rec). (~~120–130 µs/rec) — roughly two orders of magnitude. Medium 4 KiB sync/1 sat near ~21–26 MB/s; syncing every 50 climbed back to ~343–373 MB/s. fdatasync or fsync per write fell to 2 MB/sfsync and fdatasync were peers here; the real product decision is how often you pay the barrier, not which synonym you type.
Evidence path on the lab box: lab-evidence/32-fsync-vs-fdatasync/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5 on overlay/virtio (kernel 6.12). Batches: 256 B×2000 and 4 KiB×500 under /tmp. none: small p50 1.31 ms (~0.65 µs/rec, ~1.53M RPS, ~373 MB/s); fdatasync/1 259 ms; fsync/1 240 ms. Medium none 0.84 ms; fdatasync/1 75 ms; fsync/1 94 ms. every-10 / every-50 amortize (~122–373 MB/s). O_SYNC per-rec ≈ sync/1. fsync≈fdatasync on this overlay (honest). No Docker. Affiliates: 0. Evidence: lab-evidence/32-fsync-vs-fdatasync/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026