ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 71

  1. Blog
  2. /Observability & SRE

posix_fadvise SEQUENTIAL vs RANDOM: Read Hint Lab

Hands-on posix_fadvise SEQUENTIAL vs RANDOM vs NORMAL lab on localhost: cold and warm 256MiB reads after drop_caches with real measured p50 MB/s and RPS.

Aditya Challa·30 September 2026·6 min read

Lab
On this page
  1. Intro — what this post promises
  2. What posix\_fadvise changes
  3. Lab topology
  4. Arm A — cold sequential (the readahead story)
  5. Arm B — warm sequential (page cache)
  6. Arm C — random 4 KiB reads
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Methodology footnote
  11. Warm vs cold — why both tables exist
  12. Versions / environment pinned
  13. Verdict

Intro — what this post promises

posix_fadvise lets you whisper SEQUENTIAL, RANDOM, or NORMAL access patterns to the kernel. Folklore: sequential hints boost readahead; random hints stop it from wasting I/O. How much shows up on a cold 256 MiB file after drop_caches?

This is a hands-on lab with measured numbers:

  1. Cold and warm sequential 1 MiB reads under NORMAL / SEQUENTIAL / RANDOM.
  2. Cold and warm random 4 KiB reads (1500–2000 ops).
  3. Explicit drop_caches between cold runs (sudo worked on this box).
  4. Honest warm-cache results that mostly erase advice differences.

Related links:

  • mmap vs read / O_DIRECT localhost lab
  • fsync vs fdatasync localhost lab
  • sendfile vs userspace copy localhost lab
  • nice / ionice CPU and disk priority lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12, overlay/virtio). Python 3.13.5 os.posix_fadvise. File under evidence dir. Affiliates: 0. Storage is not a bare-metal NVMe paper — deltas matter more than absolute GB/s.

Verdict up front: on cold sequential reads, FADV_RANDOM cut throughput to ~989 MB/s vs NORMAL ~1685 MB/s (1.7×). FADV_SEQUENTIAL did not beat NORMAL here (1528 MB/s). On cold random 4 KiB I/O, RANDOM edged NORMAL (~1.15× RPS).


What posix_fadvise changes

AdviceIntent (simplified)
POSIX_FADV_NORMALdefault readahead policy
POSIX_FADV_SEQUENTIALexpect forward scans — more readahead
POSIX_FADV_RANDOMexpect seeks — less readahead

Hints are advisory. The kernel may ignore them; SSD firmware and page cache muddy the story further.

Related links:

  • man 2 posix_fadvise
  • mmap vs read / O_DIRECT localhost lab

Lab topology

256 MiB payload.bin (fsync'd at create)
Cold: sync; echo 3 > /proc/sys/vm/drop_caches → open → fadvise → read
Warm: prime with NORMAL full scan → fadvise → read
Seq arm: 1 MiB os.read loop to EOF
Rand arm: 1500–2000 × aligned 4 KiB pread-style lseek+read
Metric: wall p50 → MB/s or RPS + per-op p50 µs

Script: lab-evidence/38-posix-fadvise/results/run_lab.py.


Arm A — cold sequential (the readahead story)

Advicep50 wallMB/s p50
NORMAL152 ms1685
SEQUENTIAL168 ms1528 (~0.91× NORMAL)
RANDOM259 ms989 (~0.59× NORMAL)

Wrong hint hurts. Marking a full-file scan as RANDOM cost ~1.7× vs NORMAL. SEQUENTIAL was not a free win over NORMAL on this overlay/virtio stack — quote that honestly.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Arm B — warm sequential (page cache)

Advicep50MB/s
NORMAL50.4 ms5083
SEQUENTIAL48.8 ms5244
RANDOM51.6 ms4962

Once warm, advice is noise (± a few percent). Do not demo fadvise on a hot cache and claim victory.


Arm C — random 4 KiB reads

Cold (1500 ops, drop_caches each repeat):

Advicep50RPSop p50
NORMAL38.7 ms3872518.0 µs
SEQUENTIAL35.5 ms4221219.0 µs
RANDOM33.6 ms4471118.0 µs

RANDOM / NORMAL RPS ≈ 1.15× — modest, real.

Warm (2000 ops): all land ~0.5–0.6M RPS with ~1.5–1.7 µs op p50 — cache serves everything.

Related links:

  • fsync vs fdatasync localhost lab
  • sendfile vs userspace copy localhost lab

How to read these numbers

  • Match the hint to the access pattern — RANDOM on a scan was the clear footgun.
  • SEQUENTIAL ≥ NORMAL is not guaranteed on every storage stack; measure.
  • Warm cache hides advice — use drop_caches (or O_DIRECT) when claiming cold behavior.
  • Absolute GB/s here include virtio/overlay effects; use ratios in capacity debates.

Pitfalls we hit (or avoided)

  1. Benchmarking only warm reads — advice looks useless.
  2. Expecting SEQUENTIAL to always crush NORMAL — it did not on cold seq here.
  3. Forgetting drop_caches needs root — worked via sudo on this box; label if it does not on yours.
  4. Comparing random RPS without caching state — warm vs cold is an order of magnitude.
  5. Treating fadvise as a durability API — it is not fsync (see that lab).

Practical checklist

  • Log scanners / backups: prefer SEQUENTIAL or leave NORMAL; avoid RANDOM.
  • Key-value style jumps: try RANDOM; expect modest gains unless readahead was thrashing.
  • Validate with cold runs (drop_caches or direct I/O).
  • Keep warm numbers labeled so nobody ships a fake 5 GB/s “hint win.”
  • Re-measure on the real volume class (cloud block ≠ this overlay).

Related links:

  • SQLite WAL vs DELETE journal localhost lab
  • pipe vs tmpfile IPC localhost lab
  • zstd vs gzip vs lz4 compression localhost lab
  • fork COW RSS vs spawn localhost lab

Methodology footnote

File created once (256 MiB), fsyncd. Each cold trial: sync + echo 3 > /proc/sys/vm/drop_caches, short settle, open(O_RDONLY), posix_fadvise(0, size, advice), then the read loop. Warm trials prime with a NORMAL full scan first. MB/s = bytes / p50 wall; random arms also report RPS and per-op percentiles.


Warm vs cold — why both tables exist

Cold numbers answer “what happens when the kernel must fetch blocks.” Warm numbers answer “am I measuring the CPU copy loop?” Shipping only warm sequential ~5 GB/s would make every fadvise advice look identical and teach the wrong lesson. Shipping only cold without saying drop_caches ran would make the lab unreproducible on a machine where caches stay hot between trials.

On this overlay/virtio volume the absolute cold MB/s is still “fast storage” — the RANDOM vs NORMAL ~1.7× gap is the portable takeaway for sequential scanners that accidentally set the wrong hint.

If drop_caches is denied in your environment, say so and switch to a fresh file on a tmpfs-unfriendly path, or document warm-only limits like we did in older page-cache-sensitive labs.

Pair this with the mmap/O_DIRECT lab when you need to bypass page cache entirely, and with ionice when competing writers drown readahead. fadvise is a hint to the cache layer — not a scheduler class and not a durability barrier.

Versions / environment pinned

  • Python 3.13.5 with os.posix_fadvise / POSIX_FADV_*
  • 256 MiB payload.bin under the evidence directory
  • Cold path used sudo drop_caches (echo 3) successfully between trials

Bottom line for operators: treat FADV_RANDOM as a seatbelt for seek-heavy readers, not as something to sprinkle on every file descriptor. A misplaced RANDOM hint on a scanner is the failure mode this lab actually caught.

Verdict

On this box, FADV_RANDOM on a cold sequential scan dropped throughput from ~1685 MB/s (NORMAL) to ~989 MB/s. FADV_SEQUENTIAL did not beat NORMAL on that cold scan. Random 4 KiB cold I/O saw a smaller ~1.15× RPS nod to RANDOM. Hint the access you actually do — and measure cold.

Evidence path on the lab box: lab-evidence/38-posix-fadvise/results/. Affiliates: 0.

posix_fadvisefadv_sequentialfadv_randompage cachedrop_cacheslocalhost labsrelinux readahead

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; 256 MiB file; sudo drop_caches between cold runs. Cold sequential 1 MiB reads: NORMAL 1685 MB/s; SEQUENTIAL 1528 (~0.91x); RANDOM 989 (~0.59x vs NORMAL, NORMAL/RANDOM ~1.7x). Warm sequential all ~5.0-5.2 GB/s. Cold random 4 KiB x1500: RANDOM 44711 RPS vs NORMAL 38725 (~1.15x). Warm random ~0.5-0.6M RPS. Affiliates: 0. Evidence: lab-evidence/38-posix-fadvise/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What posix\_fadvise changes
  3. Lab topology
  4. Arm A — cold sequential (the readahead story)
  5. Arm B — warm sequential (page cache)
  6. Arm C — random 4 KiB reads
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Methodology footnote
  11. Warm vs cold — why both tables exist
  12. Versions / environment pinned
  13. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove