ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 67

  1. Blog

itertools.chain vs Flatten Lab

Aditya Challa·30 September 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — balanced 1 000×100 (p50)
  5. Anti-pattern scale (chain ÷ plus-loop)
  6. Consume without materializing
  7. Reading it
  8. Why plus and sum hurt
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing

Intro — what this post promises

Flattening a list-of-lists: is itertools.chain.from_iterable the right default, or is extend / a comprehension faster? This lab compares flatten strategies on Linux localhost — and times the classic foot-guns sum(lists, []) and out = out + row.

Related links:

  • itertools vs python loops localhost lab
  • nlargest vs sorted slice localhost lab
  • string concat vs join localhost lab
  • bytes vs bytearray localhost lab
  • Counter vs dict tally localhost lab
  • frozenset vs set membership localhost lab
  • attrgetter vs getattr localhost lab
  • array vs list ints localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Ops = elements written (or consumed). Affiliates: 0. Complements the general itertools lab with a flatten-only focus.

Verdict up front (balanced 1 000×100): extend ~412M/s, chain.from_iterable ~224M, nested append ~91M; sum(lists, []) ~0.96M (chain ~234× faster). Prefer extend when building a list; chain when you want an iterator.


Arms

ArmPattern
nested appenddouble for + append
extendout.extend(row)
chain.from_iterablelist(chain.from_iterable(rows))
comprehension[x for row in rows for x in row]
sum(lists, [])anti-pattern concat
out = out + rowcopy-on-each-row

Shapes: many_tiny, balanced, few_large, wide.


Lab topology

shapes: 10k×5, 1k×100, 50×5k, 5k×50
metric: p50 ops/s (elements)

Script: lab-evidence/67-itertools-chain-vs-flatten/results/run_lab.py.


Lead table — balanced 1 000×100 (p50)

Armops/sns/op
extend412,474,8962.4
chain.from_iterable224,366,6664.5
chain(*rows)228,655,0514.4
comprehension133,441,8667.5
nested append91,298,60411.0
sum(lists, [])958,2751043.5
plus-loop957,2331044.7

Anti-pattern scale (chain ÷ plus-loop)

Shapechain÷plusextend÷plus
many_tiny (10k×5)720×1079×
balanced (1k×100)234×431×
wide (5k×50)1164×1914×
few_large (50×5k)19×28×

Fewer, larger rows soften the plus/sum tax (fewer recopies) — still far behind extend/chain.


Consume without materializing

Armops/s
nested for sum67.3M
genexp sum(...)72.6M
chain.from_iterable54.6M

For pure reduce, a tight nested loop can edge chain (~1.23× here). Chain still wins on clarity when flattening into another consumer.


Reading it

  • list.extend is the throughput king for building one flat list (~4.5× nested append).
  • chain.from_iterable is the iterator default — fast enough (~2.5× nested) and lazy until list().
  • sum(lists, []) and out + row are quadratic-ish — hundreds of × slower on many rows.
  • Avoid chain(*rows) on huge outer length — argument explosion; use from_iterable.

Why plus and sum hurt

Each out = out + row allocates a new list and copies everything seen so far. With thousands of rows that is classic quadratic behavior — the same lesson as string + in a loop. sum(lists, []) is the same idea wearing a functional costume. extend amortizes growth; chain.from_iterable avoids building until you ask for a list.

If you only need to iterate once, skip materialization entirely: for x in chain.from_iterable(rows): ... (or a nested for). The consume arms show nested loops can still win a pure sum — pick clarity for pipelines, micro-optimize only on a profile.


Pitfalls

  1. Cute sum(list_of_lists, []) — slow and confusing.
  2. out = out + row in a loop — same disease as string +.
  3. Materializing with list(chain(...)) when extend suffices — extra pass.
  4. Assuming chain always beats nested — consume-only paths may differ.

When to pick what

NeedPrefer
Build one flat listextend loop
Lazy flatten / pass to consumerchain.from_iterable
One-liner listdouble comprehension
Neversum(lists, []) / repeated +

Reproduce

python3 lab-evidence/67-itertools-chain-vs-flatten/results/run_lab.py

Evidence: /workspace/lab-evidence/67-itertools-chain-vs-flatten/results/.


Closing

Extend to build; chain to iterate; never sum-concat. On this box balanced extend hit ~412M/s, chain ~224M, and sum(lists, []) trailed chain by ~234×. Flatten with the tool that matches whether you need a list or a stream.

itertools.chainfrom_iterableflattensum listsextendpythonlocalhost labsre

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. balanced: extend 412M chain 224M nested 91M; chain vs sum(lists,[]) ~234x; extend vs plus-loop ~431x. Affiliates: 0. Evidence: lab-evidence/67-itertools-chain-vs-flatten/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 72

    bytes vs bytearray: Mutate/Copy Lab

    Hands-on bytes vs bytearray lab: real ops/s for append/extend/slice/copy and when copy dominates over mutate, benchmarked on Linux localhost for SREs.

    30 Sept 2026

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 50

    signal vs threading.Event Wakeup: Localhost Lab

    Hands-on signal SIGUSR1 vs threading.Event wakeup lab: real p50 latency in microseconds, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — balanced 1 000×100 (p50)
  5. Anti-pattern scale (chain ÷ plus-loop)
  6. Consume without materializing
  7. Reading it
  8. Why plus and sum hurt
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove