ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 29

  1. Blog

statistics vs Manual mean: Stdlib Lab

Hands-on statistics vs manual mean/median/stdev lab: real ops/s for fmean, sum/len, and sorted mid versus the stdlib, measured on Linux localhost (lab).

Aditya Challa·30 September 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — mean N=10 000 (p50)
  5. Median & spread N=10 000
  6. Scale — sum/len ÷ statistics.mean
  7. Reading it
  8. Why `statistics.mean` looks “slow”
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing

Intro — what this post promises

How expensive is statistics.mean / median / stdev versus a hand sum/len, statistics.fmean, or a sorted mid? This lab times numpy-free reductions on float lists on Linux localhost.

Related links:

  • nlargest vs sorted slice localhost lab
  • bisect vs linear lookup localhost lab
  • perf_counter vs time localhost lab
  • array vs list ints localhost lab
  • Counter vs dict tally localhost lab
  • itertools chain vs flatten localhost lab
  • attrgetter vs getattr localhost lab
  • frozenset vs set membership localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. No NumPy. Correctness checks vs manual helpers: {'mean_close': True, 'fmean_close': True, 'median_close': True, 'pstdev_close': True}. Affiliates: 0. statistics.mean favors numeric care over micro-speed; fmean is the float fast path.

Verdict up front (N=10 000): sum/len ~209M elem/s vs statistics.mean ~3.4M (~62×); fmean ~24× statistics.mean. Median: manual ≈ statistics (~1.00×). Manual pstdev ~6.3× statistics.pstdev.


Arms

ArmPattern
statistics.meanstdlib exact-ish mean
statistics.fmeanfloat-optimized mean
sum(xs)/len(xs)manual
pure loop suminterpreter loop
statistics.median / manual sorted midboth sort
pstdev / stdevstdlib vs one-pass-after-mean
Welfordonline mean+sample stdev

Lab topology

N in {100, 1000, 10000, 100000}; multiple trials/size
ops = full reductions/s; also elem/s = N * trials / p50

Script: lab-evidence/69-statistics-vs-manual-mean/results/run_lab.py.


Lead table — mean N=10 000 (p50)

Armred/selem/sns/red
sum/len20,857208.6M47945
fmean8,02780.3M124582
loop sum4,17841.8M239370
statistics.mean3383.4M2961878

Median & spread N=10 000

Armred/selem/s
statistics.median9209.2M
manual median9169.2M
manual pstdev1,46814.7M
statistics.pstdev2342.3M
Welford mean+stdev1,89218.9M

Scale — sum/len ÷ statistics.mean

Nspeedup
10090×
1 00068×
10 00062×
100 00062×

fmean closes much of the gap (~24× vs statistics.mean at N=10 k) while staying in the statistics module.


Reading it

  • Hot float means — prefer statistics.fmean or sum/len; treat statistics.mean as the careful/general API.
  • Median — sorting dominates; stdlib ≈ manual here.
  • stdev/pstdev — manual two-pass beat stdlib by ~6× at N=10 k; Welford is competitive when you want one pass.
  • Correctness first — for mixed int/Fraction inputs, statistics.mean’s care may matter more than ns.

Why statistics.mean looks “slow”

On floats, statistics.mean takes a more general numeric path than fmean / a bare sum. That shows up as tens of × in a microbench — fine for reports and mixed types, noisy for a per-request float gauge. If the hot path is “average of floats I already trust,” call fmean (or sum/len) and reserve mean for clarity at API boundaries. Median’s cost is the sort either way; do not expect a helper rename to erase O(n log n).


Pitfalls

  1. Calling statistics.mean in a tight float telemetry loop — use fmean.
  2. Re-sorting for median every time on a static array — sort once.
  3. Sample vs population stdev — stdev vs pstdev (n−1 vs n).
  4. Assuming NumPy is required — stdlib covers a lot; NumPy wins on huge arrays/vectorization (out of scope).

When to pick what

NeedPrefer
Fast float meanfmean or sum/len
Mixed numeric types / docs claritystatistics.mean
Medianstatistics.median
Online mean+varianceWelford
Huge numeric arraysNumPy (elsewhere)

Reproduce

python3 lab-evidence/69-statistics-vs-manual-mean/results/run_lab.py

Evidence: /workspace/lab-evidence/69-statistics-vs-manual-mean/results/.


Closing

Use fmean for float speed; mean for generality. On this box N=10 k sum/len beat statistics.mean by ~62×, fmean by ~24×, while median stayed a wash and manual pstdev led by ~6×. Match the helper to whether you are optimizing a hot float path or writing clear stats code.

statistics.meanfmeanmedianstdevnumpy-freepythonlocalhost labsre

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. N=10k: sum/len mean 209M elem/s vs statistics.mean 3.4M (~62x); fmean ~24x statistics.mean; median ≈ parity; manual pstdev ~6.3x. Affiliates: 0. Evidence: lab-evidence/69-statistics-vs-manual-mean/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 50

    signal vs threading.Event Wakeup: Localhost Lab

    Hands-on signal SIGUSR1 vs threading.Event wakeup lab: real p50 latency in microseconds, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 76

    cmath vs math.hypot Magnitudes: Localhost Lab

    Hands-on cmath vs math.hypot magnitude ops lab: real ops/s for abs, polar, and phase, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — mean N=10 000 (p50)
  5. Median & spread N=10 000
  6. Scale — sum/len ÷ statistics.mean
  7. Reading it
  8. Why `statistics.mean` looks “slow”
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove