Plate 86
mmap vs read (+ O_DIRECT): Localhost Sequential Scan Lab
Hands-on mmap vs buffered read lab: 256 MiB p50 11.5 vs 5.4 GB/s; O_DIRECT ~0.9 GB/s warm-cache bypass. Fake endpoint-only GB/s disclosed. No Docker.
Aditya Challa5 min read
Intro — what this post promises
Need to scan a file you already opened? Folklore splits into two camps: read() into a buffer vs mmap and touch the pages. A third knob — O_DIRECT — skips the page cache on purpose. This lab measures all three on a warm box.
This is a hands-on lab with measured numbers:
- Sequential buffered
readvsmmappage-touch on 64 MiB and 256 MiB files. - Chunk-size sensitivity (4 KiB / 64 KiB / 1 MiB).
- Best-effort
O_DIRECTvia alignedlibc.read(cache bypass). - An honesty arm:
mmapthat only reads the first and last byte (fake GB/s).
Related links:
- Pipe vs tmpfile IPC localhost lab
- Unix Domain Socket vs TCP localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. Payloads under /workspace (overlay). Warm page cache — we could not drop_caches without root. O_DIRECT used a 4096-aligned buffer + libc.read. No Docker. No GPU. Affiliates: 0.
Verdict up front: at 256 MiB / 64 KiB chunks, mmap touch hit 11.5 GB/s vs buffered read 5.4 GB/s (~2.1×). O_DIRECT sat near 0.9 GB/s because it refused the warm cache. Endpoint-only mmap printed fantasy terabyte/s numbers — do not ship that graph.
What we compared
| Arm | Shape in this lab |
|---|---|
| Buffered read | open + file.read(chunk); XOR first/last byte per chunk (forces copy into userspace) |
| mmap touch | mmap.mmap + walk every 4 KiB page (fault/touch) |
| mmap endpoints only | mmap then read byte 0 and byte N−1 only (honesty / anti-pattern) |
| O_DIRECT | O_DIRECT + aligned libc.read; fall back for short tail |
Related links:
Lab topology
Arm A — 64 MiB / 64 KiB chunks (p50)
| Arm | GB/s p50 | Notes |
|---|---|---|
| Buffered read | 3.97 | Baseline userspace copy |
| mmap touch | 7.05 | ~1.78× vs buffered |
| O_DIRECT | 0.86 | direct_ok=True; cache bypass |
| mmap endpoints only | 661 | Not a scan — setup + 2 bytes |
Arm B — 256 MiB / 64 KiB chunks (p50)
| Arm | GB/s p50 | Notes |
|---|---|---|
| Buffered read | 5.37 | Larger file, still warm |
| mmap touch | 11.53 | ~2.15× vs buffered |
| O_DIRECT | 0.91 | Still ~disk/bypass class on this pass |
| mmap endpoints only | 2440 | Fantasy throughput — disclosed |
Warm-cache mmap wins here because the kernel can fault pages already resident without an extra copy into a Python bytes object on every read. Buffered read still copies. That gap is real for scan-shaped work on hot files; it is not a claim about cold HDD/NVMe.
Chunk size still matters
| Chunk | 64 MiB buffered | 64 MiB mmap touch | 64 MiB O_DIRECT |
|---|---|---|---|
| 4 KiB | 2.47 | 7.86 | 0.19 |
| 64 KiB | 3.97 | 7.05 | 0.86 |
| 1 MiB | 3.46 | 6.79 | 1.33 |
Buffered read liked mid-size chunks. O_DIRECT hated tiny reads (alignment + syscall density). mmap touch was less sensitive to the userspace chunk because the walk step was fixed at 4 KiB pages.
When buffered read or O_DIRECT still win
- Short-lived one-shot reads of small files:
readis simpler; mmap setup is overhead. - Streaming into another API that wants a
bytes/bytearrayanyway: you will copy once either way. - Isolating storage from cache effects (fio-style):
O_DIRECTis the point — “slow” on a warm box is success. - Untrusted file size / sparse horror: mmap of attacker-controlled paths needs care (not this lab’s threat model).
Pipes and tmpfiles are a different IPC question — see the pipe vs tmpfile lab.
Related links:
Pitfalls we hit (or avoided)
- Publishing mmap without page faults — endpoints-only looked 100× “faster”; we kept it as a warning.
- Calling O_DIRECT “broken” on a warm cache — it is supposed to skip cache.
- No
drop_caches— cold-cache ranking may differ; we said warm explicitly. - Python
os.read+ O_DIRECT — unaligned destination →EINVAL; alignedlibc.readfixed it. - Confusing this with
sendfile— kernel sendpath to a socket is a different lab.
Practical checklist
- Scanning a hot file in-process: try mmap (or
memoryview) and measure againstread. - Report touched bytes, not map size, when you quote GB/s.
- Use O_DIRECT when you intend to bypass cache; do not use it as a general “go faster” flag.
- Tune read chunk size; 4 KiB is often leave-performance-on-the-table.
- Do not treat loopback/overlay warm-cache numbers as cold-NVMe gospel.
Verdict
For full sequential touches on warm 64–256 MiB files, mmap beat buffered read ~1.8–2.2× (7.0 vs 4.0 GB/s at 64 MiB; 11.5 vs 5.4 GB/s at 256 MiB, 64 KiB chunks). O_DIRECT ~0.9 GB/s showed cache bypass, not failure. Endpoint-only mmap is a lie detector for your methodology — if you did not fault the pages, you did not measure the scan.
Evidence path on the lab box: lab-evidence/21-mmap-vs-read/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5. Warm page cache (no root drop_caches). 64 MiB/64 KiB: buffered 3.97 GB/s; mmap_touch 7.05; O_DIRECT 0.86 (direct_ok). 256 MiB/64 KiB: buffered 5.37; mmap_touch 11.53 (~2.15x); O_DIRECT 0.91. mmap endpoints-only 661–2440 GB/s honesty arm (not a scan). Chunk sweep 4K/64K/1M included. Affiliates: 0. Evidence: lab-evidence/21-mmap-vs-read/.
Related links
Plate 87
mmap vs read Byte-Sum Scan: Localhost Lab
Hands-on mmap vs read vs chunked sequential byte-sum scan lab: real MB/s on a generated fixture (not page-touch), measured on Linux localhost for SREs.
30 Sept 2026
Plate 32
array.array vs list[int] vs bytes Lab
A hands-on Linux localhost lab comparing memory density and numeric throughput for list[int], array.array('i'), bytearray, and memoryview.
30 Sept 2026
Plate 16
struct.pack vs to_bytes vs memoryview Lab
Benchmarking fixed-record packing with struct.pack, Struct.pack_into, int.to_bytes, and memoryview on Linux localhost.
Observability & SRE · 30 Sept 2026