Plate 92
StringIO vs list-join Builder: Localhost Lab
Aditya Challa4 min read
Intro — what this post promises
Build one large string from many chunks: io.StringIO.write, list.append + "".join, and cautionary +=. This lab reports MB/s (and builds/s) on Linux localhost.
It is not the string-concat-vs-join post (lab 49) alone — here the headline is StringIO as a file-like builder versus the classic join pattern (with += as the footgun).
Related links:
- string concat vs join localhost lab
- bytesio vs spooled tempfile localhost lab
- groupby vs manual localhost lab
- methodcaller vs getattr localhost lab
- fnmatch vs re localhost lab
- futures as completed vs wait localhost lab
- hmac compare digest localhost lab
- argparse vs sys argv localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. No Docker. Chunk = 32 ASCII chars.
Verdict up front (20 000 chunks ≈ 625 KiB): list+join ~1209.5 MB/s; StringIO ~700.7 MB/s; += ~623.4 MB/s. Prefer join for pure assembly; use StringIO when you need a file-like API (write/getvalue).
Arms
| Arm | Pattern |
|---|---|
StringIO.write then getvalue | file-like builder |
list + "".join | classic assembly |
+= in a loop | cautionary |
getvalue every write | anti-pattern (2k only) |
Lab topology
Script: lab-evidence/101-stringio-vs-list-join/results/run_lab.py.
Lead table — 20 000 chunks (p50)
| Arm | MB/s | builds/s |
|---|---|---|
| list + join | 1209.5 | 1981.6 |
| StringIO.write | 700.7 | 1148.0 |
+= loop | 623.4 | 1021.3 |
Join leads. StringIO is respectable when you want stream semantics. += trails even at 20 k chunks.
Scale sketch
| Chunks | StringIO MB/s | list+join MB/s |
|---|---|---|
| 2 000 | 832.6 | 1196.3 |
| 20 000 | 700.7 | 1209.5 |
| 100 000 | 645.8 | 1091.5 |
At 100 k chunks (~3.05 MiB), join still leads (~1091.5 vs ~645.8 MB/s).
Anti-pattern: getvalue each write
At 2 k chunks, calling getvalue() every iteration fell to ~28.8 MB/s — roughly 28.9× slower than a single final getvalue. Do not snapshot the buffer inside the hot loop.
When StringIO still wins the design
- Adapters that expect a text file-like (
write,seek,tell). - Incremental builders shared with code that already talks to files.
- Mixed write sizes where keeping a list of parts is awkward.
For “glue N known strings,” "".join(parts) remains the default.
Chunk size sensitivity
We used a fixed 32-character ASCII chunk so rates stay comparable across arms. Very tiny chunks (1–4 chars) raise call overhead and shrink absolute MB/s for everyone; huge chunks shift the bottleneck toward memory bandwidth. Re-run with your real fragment sizes before locking a design choice.
Text vs binary
StringIO is for str. For bytes pipelines use BytesIO / spooling (lab 87). Mixing encode/decode inside a hot builder will dominate any StringIO-vs-join delta measured here.
Reading it
- list+join for pure string assembly throughput.
- StringIO when the API must look like a file.
- Avoid
+=in loops of unknown length. - Avoid
getvalueper chunk.
Pitfalls
- Shipping
+=because “CPython optimizes it” — still lost here. - Measuring tiny n (noise) then extrapolating.
- Confusing BytesIO (binary) with StringIO (text) — see lab 87 for binary buffers.
- Treating lab 49’s concat story as covering StringIO — this post fills that gap.
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box. ASCII chunks only. Not Unicode-heavy normalization. Not concurrent writers.
Takeaway
At 20 k × 32-char chunks, list+join ~1209.5 MB/s beat StringIO ~700.7 MB/s and += ~623.4 MB/s. Use join for speed; StringIO for file-like builders; never getvalue each write.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. 20k×32char: list_join 1209.5 MB/s; StringIO 700.7; += 623.4. 100k: join 1091.5 vs StringIO 645.8. Not lab 49 alone. Affiliates: 0. Evidence: lab-evidence/101-stringio-vs-list-join/.
Related links
Plate 68
gc.collect Cost Empty vs Cycles: Localhost Lab
Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 58
lzma vs bz2 Compress: Localhost Lab
1 Oct 2026
Plate 57
memoryview vs bytes Slice: Localhost Lab
1 Oct 2026