ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 40

  1. Blog

functools.reduce vs for-loop: Localhost Lab

Aditya Challa·1 October 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 elem-ops/s)
  5. Reading it for SRE work
  6. Lambda tax
  7. Why reduce still exists
  8. Builtin gap
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

Aggregate sequences with functools.reduce vs a plain for loop, with sum / math.prod as reference arms. This lab reports element ops/s on Linux localhost.

Related links:

  • selectors vs select localhost lab
  • literal eval vs json localhost lab
  • cmath vs math hypot localhost lab
  • signal vs event wakeup localhost lab
  • uuid4 vs uuid1 localhost lab
  • platform vs uname localhost lab
  • shlex vs split localhost lab
  • memoryview vs bytes localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Data: ints 1..200, repeated 5000 times.

Verdict up front: builtin sum ~405562698 elem-ops/s; for-loop add ~65071639; reduce(operator.add) ~59179970; reduce(lambda) ~29566082. Product: math.prod ~29886074 vs reduce mul ~24946185.


Arms

ArmPattern
reduce(operator.add, …)fold sum
for + +=explicit sum
sum(data)C builtin
reduce(operator.mul, …)fold product
for + *=explicit product
math.prodstdlib product
reduce(lambda a,b: a+b, …)slow fold

Seven rounds, p50. Metric: (n_outer × n_elems) / p50_s.


Lab topology

n_outer=5000 · n_elems=200 · 7 rounds · p50

Script: lab-evidence/142-reduce-vs-loop/results/run_lab.py.


Lead table (p50 elem-ops/s)

Armelem-ops/s
builtin sum405562698
for-loop add65071639
reduce operator.add59179970
math.prod29886074
reduce operator.mul24946185
for-loop mul24235359
reduce lambda add29566082

Sums equal (True); prods equal (True).


Reading it for SRE work

  • Plain sums → sum() (order-of-magnitude win here).
  • Products → prefer math.prod over hand-rolled reduce.
  • reduce(operator.*) ≈ a tight loop; reduce(lambda…) paid extra call overhead (~29566082).
  • Keep reduce for non-associative folds / custom combiners — not for + on ints.

Lambda tax

reduce(lambda a, b: a + b, …) fell to ~29566082 elem-ops/s vs operator.add ~59179970. If you must fold, pass operator functions, not lambdas.


Why reduce still exists

Custom reducers (merge dicts, combine metrics structs) are clearer as reduce(combine, items, init) than a 15-line loop — just do not use it to reinvent sum. On this box the loop edged reduce for add (~65071639 vs ~59179970); builtin sum still dominated both.

Document “use sum/prod unless combiner is custom” in style guides so reviews stop bikeshedding reduce micro-optimizations.



Builtin gap

sum at ~405562698 elem-ops/s is not a fair peer to a Python-level fold — it is the reminder that CPython already optimized the common case. Reaching for reduce to “look functional” on integer sums costs readability and loses about an order of magnitude versus sum here. For products, math.prod ~29886074 still beat both reduce and the hand loop.

If a code search finds reduce(operator.add, nums, 0) in a hot metrics path, swap to sum(nums) and re-bench once — that single change usually ends the discussion.


Pitfalls

  • Teaching reduce as “the fast way” to sum.
  • Using lambda reducers in hot paths.
  • Forgetting math.prod exists (3.8+).
  • Comparing without fixing element counts.

Reproduce

python3 lab-evidence/142-reduce-vs-loop/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box, small int lists. Not numpy reductions, not parallel folds.


Prefer readable builtins first; reach for reduce only when the combiner is genuinely custom and tested.

Keep microbench scripts beside production PRs so claims stay reproducible on the same box class.

Takeaway

Use sum / math.prod for numeric aggregates (~405562698 / ~29886074 elem-ops/s). reduce(operator.add) ~59179970 tracks a for-loop; avoid lambda reduce in hot code.

functools.reducepython performancemath.prodfor loopaggregationbenchmarkingcpython

Lab evidence

What I found running this

Ran the included reduce-versus-loop benchmark on Linux localhost with Python 3.13.5. Seven rounds used 5,000 outer repetitions over 200 integers, and I compared p50 element operations per second for sum and product. Builtin sum was fastest; the for-loop edged operator.add reduce, while lambda reduce was slower. Products showed math.prod ahead of reduce and the hand loop. Results were reproducible on this box.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 58

    lzma vs bz2 Compress: Localhost Lab

    1 Oct 2026

  • Plate 34

    ast.literal_eval vs json.loads: Localhost Lab

    1 Oct 2026

  • Plate 68

    gc.collect Cost Empty vs Cycles: Localhost Lab

    Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 elem-ops/s)
  5. Reading it for SRE work
  6. Lambda tax
  7. Why reduce still exists
  8. Builtin gap
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove