Plate 40
itertools.batched vs Manual Chunking: Localhost Lab
Hands-on itertools.batched vs manual list-slice chunking: real items/s batching sequences, measured on Linux localhost today in this hands-on lab for SREs.
Aditya Challa3 min read
Intro — what this post promises
Split a sequence into fixed-size batches via itertools.batched vs manual list slices / islice chunking. This lab reports items/s on Linux localhost.
Related links:
- graphlib topo vs manual localhost lab
- functools cache vs lru localhost lab
- tomllib vs json localhost lab
- path glob vs fnmatch localhost lab
- dataclass replace vs manual localhost lab
- zoneinfo vs utc offset localhost lab
- heapq merge vs sorted localhost lab
- islice vs list slice localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. batched is stdlib since 3.12.
Verdict up front (100 k items / batch 100): batched consume ~274.18 Mitems/s; batched→list ~247.69; list slices ~242.25; islice chunks ~104.07. Prefer batched for clarity on iterables; slices fine for concrete lists.
Arms
| Arm | Pattern |
|---|---|
list(batched(seq, n)) | stdlib → list of tuples |
consume batched | no outer list |
| list comprehension slices | classic |
islice loop | iterable DIY |
| while + append slices | imperative |
Lab topology
Script: lab-evidence/120-itertools-batched-vs-chunk/results/run_lab.py.
Lead table — 100 k / batch 100 (p50)
| Arm | Mitems/s |
|---|---|
| batched consume | 274.18 |
| batched → list | 247.69 |
| list slices | 242.25 |
| while append slices | 229.33 |
| islice manual | 104.07 |
Scale sketch
| Config | batched consume | list slices |
|---|---|---|
| 10k / 256 | 296.61 | 424.81 |
| 100k / 100 | 274.18 | 242.25 |
| 100k / 1000 | 326.94 | 344.78 |
Tuple vs list batches
batched yields tuples. Wrap list(chunk) only if a callee requires lists — SQL executemany usually accepts tuples.
Remainder policy
Both styles emit a final short batch. If your protocol forbids short batches, pad or drop explicitly.
Reading it
- Any iterable / streaming source →
batched. - In-memory list → slices remain competitive.
- Avoid hand-rolled
islicechunkers unless you need custom policies. - Not the same as lab 109 (window/skip); this is tiling a sequence.
Why batch at all
Network APIs, executemany, and bulk inserts all prefer finite pages. Chunking wrong (off-by-one, dropped remainders, huge last page) causes retries and timeouts. batched makes the page loop obvious in code review.
Memory note
list(batched(...)) retains every page. For multi-million row exports, iterate for page in batched(rows, n): sink(page) without collecting. The consume arm in this lab models that path.
vs islice windows
Lab 109 measured skipping to a mid-list window. This lab tiles the whole sequence. Reusing an islice-based chunker is fine functionally but trailed batched and slices here (~104 Mitems/s vs ~240+ Mitems/s at 100 k/100).
Pitfalls
- Calling
batchedon Python ≤3.11 without a polyfill. - Materializing all batches when streaming to a sink.
- Mutating the source while iterating batches.
- Confusing batch size with worker pool size.
Reproduce
Evidence: summary.json, summary.txt.
Document batch size next to timeout and payload limits so operators can tune pages without reading the chunker implementation.
Limits
One Linux box. Integer lists in RAM. Not async or multiprocess batching.
Keep the comparison honest: measure your resource type (dummy vs FD) before optimizing.
Takeaway
At 100 k / 100, batched ~274.18 Mitems/s matches list slices ~242.25; DIY islice ~104.07 trails. Default to itertools.batched.
Lab evidence
What I found running this
Ran the batching benchmark on Linux localhost with Python 3.13.5. Measured 7 rounds and used p50 for items/s across 10k and 100k item sequences. At 100k/batch 100, batched consume measured about 274.18 Mitems/s versus 242.25 for list slices and 104.07 for islice. The result matched the expected iterable-versus-list tradeoff.
Related links
Plate 68
gc.collect Cost Empty vs Cycles: Localhost Lab
Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 16
ExitStack vs Nested with Resources: Localhost Lab
Hands-on contextlib.ExitStack vs nested with and manual close: real cycles/s for N resources, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 28
sqlite3 vs shelve Local KV: Localhost Lab
Hands-on sqlite3 vs shelve local KV store lab: real insert/get ops/s plus file sizes, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026