ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 92

  1. Blog

StringIO vs list-join Builder: Localhost Lab

Aditya Challa·30 September 2026·4 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 20 000 chunks (p50)
  5. Scale sketch
  6. Anti-pattern: getvalue each write
  7. When StringIO still wins the design
  8. Chunk size sensitivity
  9. Text vs binary
  10. Reading it
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway

Intro — what this post promises

Build one large string from many chunks: io.StringIO.write, list.append + "".join, and cautionary +=. This lab reports MB/s (and builds/s) on Linux localhost.

It is not the string-concat-vs-join post (lab 49) alone — here the headline is StringIO as a file-like builder versus the classic join pattern (with += as the footgun).

Related links:

  • string concat vs join localhost lab
  • bytesio vs spooled tempfile localhost lab
  • groupby vs manual localhost lab
  • methodcaller vs getattr localhost lab
  • fnmatch vs re localhost lab
  • futures as completed vs wait localhost lab
  • hmac compare digest localhost lab
  • argparse vs sys argv localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. No Docker. Chunk = 32 ASCII chars.

Verdict up front (20 000 chunks ≈ 625 KiB): list+join ~1209.5 MB/s; StringIO ~700.7 MB/s; += ~623.4 MB/s. Prefer join for pure assembly; use StringIO when you need a file-like API (write/getvalue).


Arms

ArmPattern
StringIO.write then getvaluefile-like builder
list + "".joinclassic assembly
+= in a loopcautionary
getvalue every writeanti-pattern (2k only)

Lab topology

chunks: 2k / 20k / 100k × 32 chars · 7 rounds · p50
metric: MB/s = out_chars / p50_s / 1024²
+= only up to 20k; getvalue-each only at 2k

Script: lab-evidence/101-stringio-vs-list-join/results/run_lab.py.


Lead table — 20 000 chunks (p50)

ArmMB/sbuilds/s
list + join1209.51981.6
StringIO.write700.71148.0
+= loop623.41021.3

Join leads. StringIO is respectable when you want stream semantics. += trails even at 20 k chunks.


Scale sketch

ChunksStringIO MB/slist+join MB/s
2 000832.61196.3
20 000700.71209.5
100 000645.81091.5

At 100 k chunks (~3.05 MiB), join still leads (~1091.5 vs ~645.8 MB/s).


Anti-pattern: getvalue each write

At 2 k chunks, calling getvalue() every iteration fell to ~28.8 MB/s — roughly 28.9× slower than a single final getvalue. Do not snapshot the buffer inside the hot loop.


When StringIO still wins the design

  • Adapters that expect a text file-like (write, seek, tell).
  • Incremental builders shared with code that already talks to files.
  • Mixed write sizes where keeping a list of parts is awkward.

For “glue N known strings,” "".join(parts) remains the default.


Chunk size sensitivity

We used a fixed 32-character ASCII chunk so rates stay comparable across arms. Very tiny chunks (1–4 chars) raise call overhead and shrink absolute MB/s for everyone; huge chunks shift the bottleneck toward memory bandwidth. Re-run with your real fragment sizes before locking a design choice.


Text vs binary

StringIO is for str. For bytes pipelines use BytesIO / spooling (lab 87). Mixing encode/decode inside a hot builder will dominate any StringIO-vs-join delta measured here.


Reading it

  • list+join for pure string assembly throughput.
  • StringIO when the API must look like a file.
  • Avoid += in loops of unknown length.
  • Avoid getvalue per chunk.

Pitfalls

  • Shipping += because “CPython optimizes it” — still lost here.
  • Measuring tiny n (noise) then extrapolating.
  • Confusing BytesIO (binary) with StringIO (text) — see lab 87 for binary buffers.
  • Treating lab 49’s concat story as covering StringIO — this post fills that gap.

Reproduce

python3 lab-evidence/101-stringio-vs-list-join/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box. ASCII chunks only. Not Unicode-heavy normalization. Not concurrent writers.


Takeaway

At 20 k × 32-char chunks, list+join ~1209.5 MB/s beat StringIO ~700.7 MB/s and += ~623.4 MB/s. Use join for speed; StringIO for file-like builders; never getvalue each write.

pythonstringiolist joinstring concatenationperformance benchmarkpython 3.13memory management

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. 20k×32char: list_join 1209.5 MB/s; StringIO 700.7; += 623.4. 100k: join 1091.5 vs StringIO 645.8. Not lab 49 alone. Affiliates: 0. Evidence: lab-evidence/101-stringio-vs-list-join/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 68

    gc.collect Cost Empty vs Cycles: Localhost Lab

    Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 58

    lzma vs bz2 Compress: Localhost Lab

    1 Oct 2026

  • Plate 57

    memoryview vs bytes Slice: Localhost Lab

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 20 000 chunks (p50)
  5. Scale sketch
  6. Anti-pattern: getvalue each write
  7. When StringIO still wins the design
  8. Chunk size sensitivity
  9. Text vs binary
  10. Reading it
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove