Plate 22
string concat vs join vs StringIO: Build Lab
A hands-on localhost lab comparing +=, list+join, preallocated join, and StringIO for building strings at real sizes.
Aditya Challa4 min read
Intro — what this post promises
Building a big string in a loop: s += chunk, parts.append + "".join, or io.StringIO? Folklore says “never +=” and “always join.” This lab builds fixed sizes from 64-byte chunks on Linux localhost and reports MB/s (p50).
Related links:
- sorted vs heapq vs bisect localhost lab
- pathlib vs os.path localhost lab
- json vs orjson vs msgpack localhost lab
- base64 vs hex encode localhost lab
- python re vs str methods localhost lab
- dataclass vs slots vs dict localhost lab
- tempfile NamedTemporaryFile localhost lab
- SHA-256 vs BLAKE2 vs xxHash localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Chunk = 64×"x". Affiliates: 0. CPython may over-allocate on += — this is not a guarantee for other runtimes.
Verdict up front: at 1 MiB, list+join ~2767 MB/s beats += ~781 (~0.28×) and StringIO ~1113 (~0.40×). At 32 MiB, join still leads (~1304 MB/s) over += (~860) and StringIO (~600). Prefer join; don’t assume StringIO is faster.
Arms
| Arm | Pattern |
|---|---|
| plus_eq | s = ""; s += chunk loop |
| list_join | parts.append(chunk); "".join(parts) |
| list_join_prealloc | pre-sized list slots then join |
| stringio | StringIO().write(chunk); getvalue() |
| bytesio / bytes join | bytes path then decode (contrast) |
Each arm must produce a string of exact target length — asserts catch under-builds.
Lab topology
Script: lab-evidence/49-string-concat-vs-join/results/run_lab.py.
Lead table — MB/s (p50)
| Size | += | list+join | prealloc join | StringIO |
|---|---|---|---|---|
| 4 KiB | 842 | 2634 | 2824 | 875 |
| 64 KiB | 854 | 2918 | 3048 | 1431 |
| 1 MiB | 781 | 2767 | 2850 | 1113 |
| 8 MiB | 992 | 2137 | 2410 | 943 |
| 32 MiB | 860 | 1304 | 1232 | 600 |
1 MiB wall times (p50): join ~0.36 ms, += ~1.28 ms, StringIO ~0.90 ms.
32 MiB: join ~24.5 ms, += ~37.2 ms, StringIO ~53.3 ms.
Prealloc list helped a little at mid sizes (~1.03× join at 1 MiB; ~1.13× at 8 MiB) and was roughly tied at 32 MiB — join itself is the win, not fancy pre-sizing.
Why += is not always a catastrophe (here)
CPython can over-allocate string storage so some += loops avoid full quadratic copies. You still see ~0.28× join at 1 MiB and ~0.66× at 32 MiB — slower, not explosive. Do not rely on that across PyPy/other versions; "".join is the portable habit.
StringIO was not the winner: write + getvalue() copy landed ~0.40× join at 1 MiB and ~0.46× at 32 MiB. Use it for file-like APIs, not as a magic faster builder.
Bytes contrast (1 MiB): BytesIO then decode ~1596 MB/s; bytes join then decode ~1262 MB/s — decode tax included; prefer staying in str or bytes end-to-end.
At 8 MiB the gap narrows some (+= ~0.46× join) but join still leads; allocator / cache effects show up in absolute MB/s — trust ratios, re-measure on your box.
Pitfalls
- Microbenching tiny strings — 4 KiB finishes in microseconds; trust 1 MiB+.
- Assuming StringIO ≈ join — measured slower here.
- += in non-CPython — may be truly quadratic.
- Building then decoding — pay once; don’t churn encodings in the loop.
- Joining with a separator by accident —
",".joinis a different problem (still usually better than loop += with commas).
When to pick what
| Need | Prefer |
|---|---|
| Build from many pieces | list + "".join |
| File-like incremental API | StringIO / real stream |
| Known byte protocol | bytearray / BytesIO / b"".join |
| One or two appends | += is fine |
Reproduce
Evidence: /workspace/lab-evidence/49-string-concat-vs-join/results/.
Closing
Join wins. On this box 1 MiB list+join ~2767 MB/s vs += ~781 vs StringIO ~1113. At 32 MiB: join ~1304, += ~860, StringIO ~600 MB/s. Default to join; keep StringIO for interfaces; treat += as acceptable only for tiny or one-off appends — and measure if the loop is hot.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5 with 64-byte chunks. Compared +=, list+join, preallocated join, StringIO, and bytes paths at 4KiB, 64KiB, 1MiB, 8MiB, and 32MiB using p50 wall time. At 1MiB list+join reached 2767 MB/s versus += 781 and StringIO 1113. At 32MiB join reached 1304 MB/s versus += 860 and StringIO 600. Preallocation helped slightly mid-size; CPython over-allocation made += slower rather than catastrophic. The measured surprise was that StringIO did not beat join. Affiliates: 0.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026