Plate 28
BytesIO vs SpooledTemporaryFile: Localhost Lab
BytesIO vs SpooledTemporaryFile growth throughput across the spool threshold, compared with NamedTemporaryFile on Linux localhost.
Aditya Challa4 min read
Intro — what this post promises
Growing a byte buffer in memory vs spilling to disk: io.BytesIO, tempfile.SpooledTemporaryFile(max_size=...), and NamedTemporaryFile. This lab reports MB/s for write+read of growing payloads on Linux localhost, under and over a 1 MiB spool threshold.
It is not a remake of the tempfile NamedTemporaryFile localhost lab (create/delete/unlink modes). Here the question is buffer growth throughput and when spooling flips to a real file.
Related links:
- tempfile namedtemporaryfile localhost lab
- mmap vs read scan localhost lab
- shutil copyfile vs manual localhost lab
- bytes vs bytearray localhost lab
- weakref vs dict cache localhost lab
- pickle vs json roundtrip localhost lab
- json dumps compact vs indent localhost lab
- decimal vs float sum localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. max_size=1 MiB, chunk 64 KiB, dir /tmp. Affiliates: 0. No Docker.
Verdict up front: ≤1 MiB — BytesIO ~16363.9 MB/s, Spooled (in-RAM) ~12426.7, Named ~1423.2. 4 MiB after spill — Spooled ~1403.6 MB/s ≈ Named ~1583.2, while BytesIO stayed ~10902.2 MB/s.
Arms
| Arm | Pattern |
|---|---|
BytesIO | pure memory grow + rewind + read |
SpooledTemporaryFile(max_size=1MiB) | memory until spill, then file |
NamedTemporaryFile | always on-disk temp file |
Lab topology
Script: lab-evidence/87-bytesio-vs-spooled-tempfile/results/run_lab.py.
Lead table (p50 MB/s)
| Size | BytesIO | Spooled | spilled? | Named |
|---|---|---|---|---|
| 256 KiB | 15523.1 | 12752.5 | False | 1527.6 |
| 1 MiB | 16363.9 | 12426.7 | False | 1423.2 |
| 4 MiB | 10902.2 | 1403.6 | True | 1583.2 |
| 16 MiB | 8424.8 | 1344.1 | True | 1382.9 |
Reading it
- Below
max_size, Spooled tracks BytesIO (same idea: memory-backed) — here ~0.76× of BytesIO at 1 MiB with a small API tax. - Past the threshold, Spooled collapses toward NamedTemporaryFile throughput (~1.3–1.6 GB/s class on this
/tmp) — the spill worked. - BytesIO remains the speed king if the whole buffer fits in RAM and you accept peak memory.
- Pick Spooled when most payloads are small but occasional large ones must not OOM the process.
Differentiation from lab 45
Lab 45 timed create/close/unlink patterns (TemporaryFile vs NamedTemporaryFile delete modes / mkstemp). This post streams growing contents and watches Spooled roll to disk at max_size.
Choosing max_size
Set max_size near your p95 payload if you want most work in RAM. Set it lower on memory-tight hosts to force earlier spill. Measure — the cliff from ~12 GB/s-class memory to ~1.4 GB/s-class disk is the operational story on this box.
Memory vs disk tradeoff
BytesIO keeps the entire buffer in the process heap — fastest, but RSS tracks payload size. NamedTemporaryFile pays syscall/fs cache costs from the first byte. SpooledTemporaryFile is the pragmatic middle: fast path for common small bodies, safe path when a rare upload exceeds max_size. On this run the spill cliff was obvious: ~12 GB/s-class memory → ~1.4 GB/s-class disk.
API tip
After writes, always seek(0) before reading — same as any file-like object. For Spooled, check whether you care that _rolled / non-BytesIO backing means the data now lives on disk (permissions, tmpdir space, unlink-on-close behavior).
Pitfalls
- Forgetting Spooled is binary mode (
w+b) for bytes. - Assuming Spooled stays fast after spill (it becomes a tempfile).
- Leaving huge BytesIO buffers alive (RSS) when Named/Spooled would bound memory.
- Benchmarking cold
/tmpvs warm page cache without saying so (we report warm sequential write/read).
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box, /tmp overlay. 64 KiB chunks. No fadvise, no tmpfs-vs-disk split beyond what /tmp is here. MB/s includes write+read round trip.
Takeaway
BytesIO ~16363.9 MB/s MiB beats disk paths by ~11.5×. Spooled stays near memory until max_size, then joins Named at ~1403.6 MB/s MiB. Use Spooled as the default “usually small, sometimes huge” buffer; use BytesIO when the cap is known and RAM is fine.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. spool_max=1MiB. 1MiB: BytesIO 16363.9 MB/s; Spooled 12426.7 (no spill); Named 1423.2. 4MiB spilled: Spooled 1403.6 ≈ Named 1583.2; BytesIO 10902.2. Not lab 45 create/delete. Affiliates: 0. Evidence: lab-evidence/87-bytesio-vs-spooled-tempfile/.
Related links
Plate 92
TemporaryFile vs NamedTemporaryFile Delete Modes Lab
Hands-on tempfile modes lab: TemporaryFile vs NamedTemporaryFile delete True/False vs mkstemp unlink patterns with real measured ops/s on localhost box.
Observability & SRE · 30 Sept 2026
Plate 17
platform vs os.uname Inventory: Localhost Lab
Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 75
uuid.uuid4 vs uuid.uuid1: Localhost Lab
Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026