ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 87

  1. Blog

mmap vs read Byte-Sum Scan: Localhost Lab

Hands-on mmap vs read vs chunked sequential byte-sum scan lab: real MB/s on a generated fixture (not page-touch), measured on Linux localhost for SREs.

Aditya Challa·30 September 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 64 MiB (p50)
  5. 16 MiB check
  6. Reading it
  7. Differentiation from lab 21
  8. Memory pressure angle
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

How fast can Python sum every byte of a file — open().read(), chunked read, or mmap? This lab reports MB/s for a full sequential byte-sum on generated fixtures on Linux localhost.

It is not a remake of the mmap vs read O_DIRECT localhost lab, which measured page-touch / XOR-style scan throughput in GB/s plus an O_DIRECT arm. Here the metric is deliberately Python sum(memoryview(...)) over every byte — a compute-heavy sequential scan where the interpreter dominates, so mmap vs buffered read gaps shrink and chunk size matters differently.

Related links:

  • mmap vs read odirect localhost lab
  • processpoolexecutor vs sequential localhost lab
  • csv reader vs split localhost lab
  • pickle vs json roundtrip localhost lab
  • bytes vs bytearray localhost lab
  • hashlib md5 vs blake2b localhost lab
  • shutil copyfile vs manual localhost lab
  • glob vs rglob vs walk localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Warm page cache (no root drop_caches). Fixtures fixture_16MiB.bin / fixture_64MiB.bin under evidence. Affiliates: 0. No O_DIRECT. No Docker.

Verdict up front (64 MiB byte-sum): mmap ~197.4 MB/s vs read-all ~179.4 MB/s (~1.1×); chunked 64 KiB/1 MiB ~191.3–192.4 MB/s; 4 KiB chunks slower (~181.4 MB/s). All arms live near ~180–200 MB/s because summing bytes in Python is the bottleneck — not disk.


Arms

ArmPattern
read_all_sumf.read() then sum(memoryview(data))
mmap_summmap then sum(memoryview(mm))
chunked_read_*loop read(chunk) + sum each chunk
mmap_chunked_65536mmap + sum 64 KiB windows

Lab topology

fixtures: 16 MiB / 64 MiB · rounds=5 · p50 wall
chunks: 4 KiB / 64 KiB / 1 MiB
metric: MB/s = size_MiB / p50_s
warm page cache

Script: lab-evidence/81-mmap-vs-read-scan/results/run_lab.py.


Lead table — 64 MiB (p50)

Armp50 sMB/s
read_all_sum0.3567179.4
mmap_sum0.3242197.4
chunked 4 KiB0.3528181.4
chunked 64 KiB0.3346191.3
chunked 1 MiB0.3326192.4
mmap_chunked 64 KiB0.326196.3

16 MiB check

ArmMB/s
read_all184.6
mmap190.2
chunked 64 KiB189.2
chunked 1 MiB190.2
mmap_chunked 64 KiB196.2

Same band — confirms the interpreter/sum path, not fixture quirks.


Reading it

  • Byte-sum in Python is CPU-bound at ~180–200 MB/s here; mmap’s classic GB/s win from lab 21 does not transfer when every byte is reduced in the interpreter.
  • mmap still edged read-all by ~1.1× at 64 MiB (slightly less peak RSS from avoiding a second full buffer, plus map overhead differences).
  • Chunk size: 4 KiB loses to call overhead; 64 KiB–1 MiB cluster with mmap.
  • Prefer mmap or large-chunked read when you must scan without holding two full copies — but for pure Python reductions, rewrite hot loops in C/array/numpy before obsessing over mmap.

Differentiation from lab 21

Lab 21’s mmap touch walked pages and reported multi-GB/s; O_DIRECT exposed cache bypass. This lab’s sum(memoryview) forces a visit to every byte in Python, so throughput collapses into the ~0.2 GB/s class. Use lab 21 when the question is “how fast can I fault/touch pages?” Use this post when the question is “how fast can I reduce file bytes in Python?”


Memory pressure angle

read_all allocates a contiguous bytes object equal to file size plus the temporary work of summing. mmap maps pages and can avoid duplicating the whole file in the process heap (OS still caches pages). On memory-tight hosts, mmap/chunked paths matter even when MB/s looks similar.


Pitfalls

  • Comparing mmap to a page-touch microbench and claiming the same GB/s for sum/hash loops.
  • Tiny chunks (4 KiB) drowning in syscall/Python overhead.
  • Leaving a memoryview exported while closing mmap (BufferError) — release views first.
  • Cold-cache first pass vs warm-cache (we report warm; cold will be slower and more disk-bound).

Reproduce

python3 lab-evidence/81-mmap-vs-read-scan/results/run_lab.py

Evidence includes fixture_16MiB.bin, fixture_64MiB.bin, summary.json.


Limits

Warm cache only. Overlay//workspace storage. No O_DIRECT, no madvise, no multi-threaded scan. Numbers are for Python byte-sum, not memcpy bandwidth.


Takeaway

For a sequential Python byte-sum, mmap (~197.4 MB/s) and large-chunked read (~192.4 MB/s) beat tiny chunks and slightly beat read-all (~179.4 MB/s) at 64 MiB — but all sit in the same ~180–200 MB/s band. The big win is usually less peak memory and sane chunking, not a magical mmap GB/s number from a different kind of scan.

mmap vs readbyte-sum scanchunked readmemoryviewsequential scanlocalhost labsrepython mmap

Lab evidence

What I found running this

Ran the Linux localhost byte-sum scan lab on 1 Oct 2026 IST with Python 3.13.5 and warm page cache. On the 64 MiB fixture, mmap measured 197.4 MB/s versus read_all 179.4 MB/s; 64 KiB and 1 MiB chunked reads reached 191.3 and 192.4 MB/s, while 4 KiB reached 181.4 MB/s. Affiliates: 0.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 86

    mmap vs read (+ O_DIRECT): Localhost Sequential Scan Lab

    Hands-on mmap vs buffered read lab: 256 MiB p50 11.5 vs 5.4 GB/s; O_DIRECT ~0.9 GB/s warm-cache bypass. Fake endpoint-only GB/s disclosed. No Docker.

    30 Sept 2026

  • Plate 32

    array.array vs list[int] vs bytes Lab

    A hands-on Linux localhost lab comparing memory density and numeric throughput for list[int], array.array('i'), bytearray, and memoryview.

    30 Sept 2026

  • Plate 16

    struct.pack vs to_bytes vs memoryview Lab

    Benchmarking fixed-record packing with struct.pack, Struct.pack_into, int.to_bytes, and memoryview on Linux localhost.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 64 MiB (p50)
  5. 16 MiB check
  6. Reading it
  7. Differentiation from lab 21
  8. Memory pressure angle
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove