ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 28

  1. Blog

BytesIO vs SpooledTemporaryFile: Localhost Lab

BytesIO vs SpooledTemporaryFile growth throughput across the spool threshold, compared with NamedTemporaryFile on Linux localhost.

Aditya Challa·30 September 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 MB/s)
  5. Reading it
  6. Differentiation from lab 45
  7. Choosing max\_size
  8. Memory vs disk tradeoff
  9. API tip
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway

Intro — what this post promises

Growing a byte buffer in memory vs spilling to disk: io.BytesIO, tempfile.SpooledTemporaryFile(max_size=...), and NamedTemporaryFile. This lab reports MB/s for write+read of growing payloads on Linux localhost, under and over a 1 MiB spool threshold.

It is not a remake of the tempfile NamedTemporaryFile localhost lab (create/delete/unlink modes). Here the question is buffer growth throughput and when spooling flips to a real file.

Related links:

  • tempfile namedtemporaryfile localhost lab
  • mmap vs read scan localhost lab
  • shutil copyfile vs manual localhost lab
  • bytes vs bytearray localhost lab
  • weakref vs dict cache localhost lab
  • pickle vs json roundtrip localhost lab
  • json dumps compact vs indent localhost lab
  • decimal vs float sum localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. max_size=1 MiB, chunk 64 KiB, dir /tmp. Affiliates: 0. No Docker.

Verdict up front: ≤1 MiB — BytesIO ~16363.9 MB/s, Spooled (in-RAM) ~12426.7, Named ~1423.2. 4 MiB after spill — Spooled ~1403.6 MB/s ≈ Named ~1583.2, while BytesIO stayed ~10902.2 MB/s.


Arms

ArmPattern
BytesIOpure memory grow + rewind + read
SpooledTemporaryFile(max_size=1MiB)memory until spill, then file
NamedTemporaryFilealways on-disk temp file

Lab topology

sizes: 256 KiB / 1 / 4 / 16 MiB · chunk 64 KiB · 7 rounds · p50
metric: MB/s = size_MiB / p50_wall (write+read)
spool_max_size = 1048576 bytes

Script: lab-evidence/87-bytesio-vs-spooled-tempfile/results/run_lab.py.


Lead table (p50 MB/s)

SizeBytesIOSpooledspilled?Named
256 KiB15523.112752.5False1527.6
1 MiB16363.912426.7False1423.2
4 MiB10902.21403.6True1583.2
16 MiB8424.81344.1True1382.9

Reading it

  • Below max_size, Spooled tracks BytesIO (same idea: memory-backed) — here ~0.76× of BytesIO at 1 MiB with a small API tax.
  • Past the threshold, Spooled collapses toward NamedTemporaryFile throughput (~1.3–1.6 GB/s class on this /tmp) — the spill worked.
  • BytesIO remains the speed king if the whole buffer fits in RAM and you accept peak memory.
  • Pick Spooled when most payloads are small but occasional large ones must not OOM the process.

Differentiation from lab 45

Lab 45 timed create/close/unlink patterns (TemporaryFile vs NamedTemporaryFile delete modes / mkstemp). This post streams growing contents and watches Spooled roll to disk at max_size.


Choosing max_size

Set max_size near your p95 payload if you want most work in RAM. Set it lower on memory-tight hosts to force earlier spill. Measure — the cliff from ~12 GB/s-class memory to ~1.4 GB/s-class disk is the operational story on this box.


Memory vs disk tradeoff

BytesIO keeps the entire buffer in the process heap — fastest, but RSS tracks payload size. NamedTemporaryFile pays syscall/fs cache costs from the first byte. SpooledTemporaryFile is the pragmatic middle: fast path for common small bodies, safe path when a rare upload exceeds max_size. On this run the spill cliff was obvious: ~12 GB/s-class memory → ~1.4 GB/s-class disk.


API tip

After writes, always seek(0) before reading — same as any file-like object. For Spooled, check whether you care that _rolled / non-BytesIO backing means the data now lives on disk (permissions, tmpdir space, unlink-on-close behavior).


Pitfalls

  • Forgetting Spooled is binary mode (w+b) for bytes.
  • Assuming Spooled stays fast after spill (it becomes a tempfile).
  • Leaving huge BytesIO buffers alive (RSS) when Named/Spooled would bound memory.
  • Benchmarking cold /tmp vs warm page cache without saying so (we report warm sequential write/read).

Reproduce

python3 lab-evidence/87-bytesio-vs-spooled-tempfile/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box, /tmp overlay. 64 KiB chunks. No fadvise, no tmpfs-vs-disk split beyond what /tmp is here. MB/s includes write+read round trip.


Takeaway

BytesIO ~16363.9 MB/s  MiB beats disk paths by ~11.5×. Spooled stays near memory until max_size, then joins Named at ~1403.6 MB/s  MiB. Use Spooled as the default “usually small, sometimes huge” buffer; use BytesIO when the cap is known and RAM is fine.

bytesiospooledtemporaryfilenamedtemporaryfilebuffer growthmax_size spoollocalhost labsrepython tempfile

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. spool_max=1MiB. 1MiB: BytesIO 16363.9 MB/s; Spooled 12426.7 (no spill); Named 1423.2. 4MiB spilled: Spooled 1403.6 ≈ Named 1583.2; BytesIO 10902.2. Not lab 45 create/delete. Affiliates: 0. Evidence: lab-evidence/87-bytesio-vs-spooled-tempfile/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 92

    TemporaryFile vs NamedTemporaryFile Delete Modes Lab

    Hands-on tempfile modes lab: TemporaryFile vs NamedTemporaryFile delete True/False vs mkstemp unlink patterns with real measured ops/s on localhost box.

    Observability & SRE · 30 Sept 2026

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 75

    uuid.uuid4 vs uuid.uuid1: Localhost Lab

    Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 MB/s)
  5. Reading it
  6. Differentiation from lab 45
  7. Choosing max\_size
  8. Memory vs disk tradeoff
  9. API tip
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove