ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 86

  1. Blog

mmap vs read (+ O_DIRECT): Localhost Sequential Scan Lab

Hands-on mmap vs buffered read lab: 256 MiB p50 11.5 vs 5.4 GB/s; O_DIRECT ~0.9 GB/s warm-cache bypass. Fake endpoint-only GB/s disclosed. No Docker.

Aditya Challa·30 September 2026·5 min read

Summary
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — 64 MiB / 64 KiB chunks (p50)
  5. Arm B — 256 MiB / 64 KiB chunks (p50)
  6. Chunk size still matters
  7. When buffered read or O\_DIRECT still win
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict

Intro — what this post promises

Need to scan a file you already opened? Folklore splits into two camps: read() into a buffer vs mmap and touch the pages. A third knob — O_DIRECT — skips the page cache on purpose. This lab measures all three on a warm box.

This is a hands-on lab with measured numbers:

  1. Sequential buffered read vs mmap page-touch on 64 MiB and 256 MiB files.
  2. Chunk-size sensitivity (4 KiB / 64 KiB / 1 MiB).
  3. Best-effort O_DIRECT via aligned libc.read (cache bypass).
  4. An honesty arm: mmap that only reads the first and last byte (fake GB/s).

Related links:

  • Pipe vs tmpfile IPC localhost lab
  • Unix Domain Socket vs TCP localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. Payloads under /workspace (overlay). Warm page cache — we could not drop_caches without root. O_DIRECT used a 4096-aligned buffer + libc.read. No Docker. No GPU. Affiliates: 0.

Verdict up front: at 256 MiB / 64 KiB chunks, mmap touch hit 11.5 GB/s vs buffered read 5.4 GB/s (~2.1×). O_DIRECT sat near 0.9 GB/s because it refused the warm cache. Endpoint-only mmap printed fantasy terabyte/s numbers — do not ship that graph.


What we compared

ArmShape in this lab
Buffered readopen + file.read(chunk); XOR first/last byte per chunk (forces copy into userspace)
mmap touchmmap.mmap + walk every 4 KiB page (fault/touch)
mmap endpoints onlymmap then read byte 0 and byte N−1 only (honesty / anti-pattern)
O_DIRECTO_DIRECT + aligned libc.read; fall back for short tail

Related links:

  • man 2 mmap
  • man 2 open — O_DIRECT

Lab topology

payload_64mb.bin / payload_256mb.bin
rounds=5 each · chunks 4 KiB, 64 KiB, 1 MiB
Arms: read_buffered | mmap_touch | mmap_endpoints_only | read_odirect
Metric: wall GB/s (p50 / mean) for a full sequential pass

Arm A — 64 MiB / 64 KiB chunks (p50)

ArmGB/s p50Notes
Buffered read3.97Baseline userspace copy
mmap touch7.05~1.78× vs buffered
O_DIRECT0.86direct_ok=True; cache bypass
mmap endpoints only661Not a scan — setup + 2 bytes

Arm B — 256 MiB / 64 KiB chunks (p50)

ArmGB/s p50Notes
Buffered read5.37Larger file, still warm
mmap touch11.53~2.15× vs buffered
O_DIRECT0.91Still ~disk/bypass class on this pass
mmap endpoints only2440Fantasy throughput — disclosed

Warm-cache mmap wins here because the kernel can fault pages already resident without an extra copy into a Python bytes object on every read. Buffered read still copies. That gap is real for scan-shaped work on hot files; it is not a claim about cold HDD/NVMe.


Chunk size still matters

Chunk64 MiB buffered64 MiB mmap touch64 MiB O_DIRECT
4 KiB2.477.860.19
64 KiB3.977.050.86
1 MiB3.466.791.33

Buffered read liked mid-size chunks. O_DIRECT hated tiny reads (alignment + syscall density). mmap touch was less sensitive to the userspace chunk because the walk step was fixed at 4 KiB pages.


When buffered read or O_DIRECT still win

  • Short-lived one-shot reads of small files: read is simpler; mmap setup is overhead.
  • Streaming into another API that wants a bytes/bytearray anyway: you will copy once either way.
  • Isolating storage from cache effects (fio-style): O_DIRECT is the point — “slow” on a warm box is success.
  • Untrusted file size / sparse horror: mmap of attacker-controlled paths needs care (not this lab’s threat model).

Pipes and tmpfiles are a different IPC question — see the pipe vs tmpfile lab.

Related links:

  • Pipe vs tmpfile IPC localhost lab

Pitfalls we hit (or avoided)

  1. Publishing mmap without page faults — endpoints-only looked 100× “faster”; we kept it as a warning.
  2. Calling O_DIRECT “broken” on a warm cache — it is supposed to skip cache.
  3. No drop_caches — cold-cache ranking may differ; we said warm explicitly.
  4. Python os.read + O_DIRECT — unaligned destination → EINVAL; aligned libc.read fixed it.
  5. Confusing this with sendfile — kernel sendpath to a socket is a different lab.

Practical checklist

  • Scanning a hot file in-process: try mmap (or memoryview) and measure against read.
  • Report touched bytes, not map size, when you quote GB/s.
  • Use O_DIRECT when you intend to bypass cache; do not use it as a general “go faster” flag.
  • Tune read chunk size; 4 KiB is often leave-performance-on-the-table.
  • Do not treat loopback/overlay warm-cache numbers as cold-NVMe gospel.

Verdict

For full sequential touches on warm 64–256 MiB files, mmap beat buffered read ~1.8–2.2× (7.0 vs 4.0 GB/s at 64 MiB; 11.5 vs 5.4 GB/s at 256 MiB, 64 KiB chunks). O_DIRECT ~0.9 GB/s showed cache bypass, not failure. Endpoint-only mmap is a lie detector for your methodology — if you did not fault the pages, you did not measure the scan.

Evidence path on the lab box: lab-evidence/21-mmap-vs-read/results/. Affiliates: 0.

mmap vs reado_directpage cachesequential scanlinux performancelocalhost labsrememoryview

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5. Warm page cache (no root drop_caches). 64 MiB/64 KiB: buffered 3.97 GB/s; mmap_touch 7.05; O_DIRECT 0.86 (direct_ok). 256 MiB/64 KiB: buffered 5.37; mmap_touch 11.53 (~2.15x); O_DIRECT 0.91. mmap endpoints-only 661–2440 GB/s honesty arm (not a scan). Chunk sweep 4K/64K/1M included. Affiliates: 0. Evidence: lab-evidence/21-mmap-vs-read/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 87

    mmap vs read Byte-Sum Scan: Localhost Lab

    Hands-on mmap vs read vs chunked sequential byte-sum scan lab: real MB/s on a generated fixture (not page-touch), measured on Linux localhost for SREs.

    30 Sept 2026

  • Plate 32

    array.array vs list[int] vs bytes Lab

    A hands-on Linux localhost lab comparing memory density and numeric throughput for list[int], array.array('i'), bytearray, and memoryview.

    30 Sept 2026

  • Plate 16

    struct.pack vs to_bytes vs memoryview Lab

    Benchmarking fixed-record packing with struct.pack, Struct.pack_into, int.to_bytes, and memoryview on Linux localhost.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — 64 MiB / 64 KiB chunks (p50)
  5. Arm B — 256 MiB / 64 KiB chunks (p50)
  6. Chunk size still matters
  7. When buffered read or O\_DIRECT still win
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove