ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 60

  1. Blog
  2. /Observability & SRE

fsync vs fdatasync vs none: Durable Write Cost Lab

Hands-on fsync vs fdatasync vs none: per-write and every-N sync on 256B and 4KiB records. Real p50/p95 show about 100-200x durability tax on localhost

Aditya Challa·30 September 2026·7 min read

Lab
On this page
  1. Intro — what this post promises
  2. What fsync / fdatasync / O\_SYNC actually change
  3. Lab topology
  4. Arm A — small records (256 B × 2000)
  5. Arm B — medium records (4 KiB × 500)
  6. Arm C — amortize: every 10 / every 50
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict

Intro — what this post promises

Calling write() is not the same as durable data. The kernel is happy to park your bytes in the page cache and return. Crash (or power loss) later and those “writes” never happened.

This is a hands-on lab with measured numbers:

  1. Wall time for batches of small (256 B) and medium (4 KiB) records with no sync.
  2. Same batches with fdatasync per write and fsync per write.
  3. Amortization: sync every 10 / every 50 writes.
  4. Optional O_SYNC open — durability on every write(2).
  5. Honest limits on an overlay/virtio lab box (not a battery-backed WAL appliance).

Related links:

  • mmap vs read / O_DIRECT localhost lab
  • sendfile vs userspace copy localhost lab
  • pipe vs tmpfile IPC localhost lab
  • nice / ionice CPU and disk priority lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 writing under /tmp on overlay atop virtio. No Docker. No public bind. Affiliates: 0. This measures sync cost, not crash recovery or fsck after echo b > /proc/sysrq-trigger.

Verdict up front: none → any per-write sync was the cliff (~100–200×). On this overlay, fsync ≈ fdatasync (within noise). Syncing every 50 medium records bought back hundreds of MB/s.


What fsync / fdatasync / O_SYNC actually change

KnobWhat it waits for (simplified)Typical use
none (buffered write)bytes in page cachetemp files, best-effort logs
fdatasyncfile data (and needed metadata for that data) on storageWAL / record append without full inode flush folklore
fsyncdata and metadatarename-over, directory durability stories
O_SYNC / O_DSYNCeach write behaves like write+syncsimple but brutal APIs

Textbook claim: fdatasync is cheaper than fsync because it can skip some metadata. On this lab FS the gap was tiny — report that honestly instead of pasting a textbook 2×.

Related links:

  • man 2 fsync
  • mmap vs read / O_DIRECT localhost lab

Lab topology

Python 3.13 · /tmp on overlay/virtio
Arms: none | fdatasync every N | fsync every N | O_SYNC open
Shapes: 256 B × 2000 records · 4 KiB × 500 records
Metric: batch wall p50/p95 → per-record µs, RPS, MB/s
Warmup batches discarded; repeats 4–8 depending on arm

Script: lab-evidence/32-fsync-vs-fdatasync/results/run_lab.py. Each arm creates/truncates a file, writes N fixed payloads, applies the sync policy, closes, records perf_counter wall time for the whole batch.

RPS / MB/s below are from p50 batch wall (ops and bytes ÷ p50 seconds) so a slow repeat does not silently dilute the headline.


Arm A — small records (256 B × 2000)

Armp50 batchp95per-rec p50RPS (p50)MB/s (p50)
none1.31 ms2.720.65 µs~1.53M~373
fdatasync / 1259 ms268130 µs~7.7k~1.9
fsync / 1240 ms261120 µs~8.3k~2.0
O_SYNC (500 rec)60.4 ms62.7121 µs~8.3k~2.0

none → fdatasync/1 ≈ 198× on batch p50. That is the durability tax in one table.

fsync/1 was not slower than fdatasync/1 here (slightly faster on this run). Treat them as tied on overlay, not as a portable “always use fdatasync” win.

Stability×3 for none p50 (ms): 1.18 / 1.26 / 1.18. For fdatasync/1: 246 / 234 / 233.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Arm B — medium records (4 KiB × 500)

Armp50 batchp95per-rec p50RPS (p50)MB/s (p50)
none0.84 ms1.491.7 µs~593k~2317
fdatasync / 175.1 ms88.7150 µs~6.7k~26
fsync / 193.5 ms99.3187 µs~5.3k~21
O_SYNC (200 rec)35.3 ms37.6176 µs~5.7k~22

Medium none looks like memory copy (multi‑GB/s class) because it mostly dirties cache. Per-write sync drops you into the ~20–26 MB/s band on this box — still faster than the 256 B sync arms in MB/s (larger payload per barrier) but nowhere near “disk sequential” marketing numbers.

Stability×3 for med fsync/1 p50 (ms): 85.0 / 68.6 / 77.5 (noisy — quote bands, not fake precision).


Arm C — amortize: every 10 / every 50

Same 4 KiB × 500 shape; sync only every N records (plus a final sync if the last group is short).

Armp50per-rec p50MB/s (p50)vs sync/1
fdatasync / 1016.0 ms31.9 µs~122~4.7× vs /1
fsync / 1012.8 ms25.6 µs~152~7.3×
fdatasync / 505.70 ms11.4 µs~343~13×
fsync / 505.24 ms10.5 µs~373~18×

Group commit is not a metaphor: every-50 recovered into the hundreds of MB/s while still issuing durability barriers. Databases and queues do this for a reason.

Related links:

  • pipe vs tmpfile IPC localhost lab
  • flock contention exclusive vs shared lab

How to read these numbers

  • none ≠ durable. Fast path is page cache. Fine for scratch; wrong for “acked to disk.”
  • Per-write sync turns throughput into barrier rate. Small records hurt more in MB/s.
  • fsync vs fdatasync: textbook gap did not show on this overlay run — do not ship a fake 3× claim from this post.
  • O_SYNC: per-record cost matched sync/1 (~120–180 µs). Convenient API; same economic cliff.
  • Every-N is the practical dial between RPO fantasy and RPS reality.

Pitfalls we hit (or avoided)

  1. Reporting sum-of-repeats RPS — early harness bug understated RPS by ×repeats; numbers above use p50 batch wall.
  2. Assuming fdatasync always wins — not on this FS snapshot.
  3. Calling none “disk throughput” — it is not; see mmap / O_DIRECT labs for device-ish paths.
  4. One p50 without p95 — sync arms were tighter relative to mean; still show both.
  5. Claiming crash safety — we timed syscalls; we did not pull the plug.

Practical checklist

  • Decide the durability contract (lost last N records OK? or not?).
  • If you need sync, batch (every N / group commit / fdatasync on a WAL segment) before blaming the SSD.
  • Benchmark your FS — overlay ≠ XFS on NVMe ≠ cloud block with flaky flush.
  • Pair with p95 and a clear record size; do not quote only MB/s from none.
  • Keep evidence next to any “we fsync every message” capacity plan.

Related links:

  • stdout buffering line vs full localhost lab
  • context-switch pipe ping-pong localhost lab
  • sendfile vs userspace copy localhost lab

Verdict

On this box, buffered none wrote 256 B records at 373 MB/s~~ (0.65 µs/rec). fdatasync or fsync per write fell to 2 MB/s (~~120–130 µs/rec) — roughly two orders of magnitude. Medium 4 KiB sync/1 sat near ~21–26 MB/s; syncing every 50 climbed back to ~343–373 MB/s. fsync and fdatasync were peers here; the real product decision is how often you pay the barrier, not which synonym you type.

Evidence path on the lab box: lab-evidence/32-fsync-vs-fdatasync/results/. Affiliates: 0.

fsync vs fdatasyncdurable writeso_syncwrite amplificationlinux durabilitylocalhost labsrewal sync

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5 on overlay/virtio (kernel 6.12). Batches: 256 B×2000 and 4 KiB×500 under /tmp. none: small p50 1.31 ms (~0.65 µs/rec, ~1.53M RPS, ~373 MB/s); fdatasync/1 259 ms; fsync/1 240 ms. Medium none 0.84 ms; fdatasync/1 75 ms; fsync/1 94 ms. every-10 / every-50 amortize (~122–373 MB/s). O_SYNC per-rec ≈ sync/1. fsync≈fdatasync on this overlay (honest). No Docker. Affiliates: 0. Evidence: lab-evidence/32-fsync-vs-fdatasync/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What fsync / fdatasync / O\_SYNC actually change
  3. Lab topology
  4. Arm A — small records (256 B × 2000)
  5. Arm B — medium records (4 KiB × 500)
  6. Arm C — amortize: every 10 / every 50
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove