ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 94

  1. Blog

statistics.quantiles vs Manual Percentile: Localhost Lab

Aditya Challa·1 October 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — n=10 000 (p50 ops/s)
  5. Scale sketch (exclusive quantiles vs manual sort)
  6. Reading it
  7. Method names matter
  8. Sort once
  9. Streaming / online note
  10. Comparison to mean lab
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway

Intro — what this post promises

Compute quartiles / percentiles on a float list via statistics.quantiles vs a sorted-index manual method. This lab reports ops/s on Linux localhost.

It is not statistics-vs-manual-mean (lab 69) — that was means; this is quantiles/percentiles.

Related links:

  • statistics vs manual mean localhost lab
  • graphlib topo vs manual localhost lab
  • functools cache vs lru localhost lab
  • tomllib vs json localhost lab
  • path glob vs fnmatch localhost lab
  • dataclass replace vs manual localhost lab
  • zoneinfo vs utc offset localhost lab
  • heapq merge vs sorted localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Uniform random floats.

Verdict up front (n=10 000): quantiles exclusive ~903.9 ops/s; manual sort+index ~904.3; quantiles on already-sorted input ~11596.9; index-only on pre-sorted ~876423.3. Sorting dominates; prefer stdlib for method semantics.


Arms

ArmPattern
statistics.quantiles(..., method="exclusive")stdlib
method="inclusive"alternate cuts
quantiles on pre-sorted liststill sorts inside
manual sort + linear indexteaching recipe
index-only if pre-sortedskip sort tax
nearest-rank 25/50/75simple percentiles

Lab topology

n in {1000, 10000, 100000} · 7 rounds · p50
metric: ops/s = 1 / p50_s (one quartile/percentile pass)

Script: lab-evidence/116-statistics-quantiles-vs-manual/results/run_lab.py.


Lead table — n=10 000 (p50 ops/s)

Armops/s
manual pre-sorted index-only876423.3
quantiles (pre-sorted input)11596.9
manual sort + linear904.3
statistics exclusive903.9
nearest 25/50/75807.2
statistics inclusive759.3

Scale sketch (exclusive quantiles vs manual sort)

nstatistics exclusivemanual sort+index
1 00010900.712343.4
10 000903.9904.3
100 00048.964.2

Throughput collapses with n log n sort cost — expected.


Reading it

  • Unsorted data → statistics.quantiles (clear methods, maintained).
  • Already sorted stream → index math can win big (~876423.3 ops/s here) but document the rule.
  • Do not confuse with mean benches (lab 69).
  • Inclusive vs exclusive change cut points — pick one and stick to it in dashboards.

Method names matter

exclusive vs inclusive exist because “quartile” is not one universal formula. Shipping a blog-post index recipe into finance/report pipelines without naming the method causes silent dashboard drift. Prefer documenting statistics.quantiles(..., method=...) in the runbook.


Sort once

If you need p50, p90, p99, sort (or maintain a sorted structure) once, then index. Repeated full sorts per percentile were not the happy path on this box at n=100k (~48.9 ops/s exclusive).


Streaming / online note

These arms assume an in-memory list. True online percentiles (t-digest, P², HDR histograms) are a different design space — do not paste quantiles(list(all_samples)) into a hot telemetry path that retains every float forever. Batch windows of fixed size, then call stdlib once per window.


Comparison to mean lab

Lab 69’s mean/fsum story is O(n) without a full sort. Quartiles pay for ordering. If a dashboard only needs central tendency, do not upgrade it to quantiles “for free” — you buy n log n unless the array is already sorted for another reason.


Pitfalls

  • DIY percentiles that disagree with Excel/pandas/quantiles method.
  • Assuming pre-sorted input skips work inside statistics.quantiles (it still sorts).
  • Computing many one-off percentiles each with a full sort — sort once.
  • Floating ties and tiny samples (unstable cuts).

Reproduce

python3 lab-evidence/116-statistics-quantiles-vs-manual/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box. Uniform floats only. Manual arms are illustrative — not bit-identical to every method=. Not numpy/pandas.


Takeaway

At 10 k floats, quantiles ~903.9 ops/s ≈ manual sort+index ~904.3. Prefer stdlib quantiles for semantics; reuse a sorted buffer if you need many cuts.

statistics.quantilespercentilequartilepython statisticslocalhost labsreops/s

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. n=10k: quantiles exclusive 903.9 ops/s; manual sort 904.3; pre-sorted index-only 876423.3. Not lab 69 mean. Affiliates: 0. Evidence: lab-evidence/116-statistics-quantiles-vs-manual/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 75

    uuid.uuid4 vs uuid.uuid1: Localhost Lab

    Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 76

    cmath vs math.hypot Magnitudes: Localhost Lab

    Hands-on cmath vs math.hypot magnitude ops lab: real ops/s for abs, polar, and phase, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — n=10 000 (p50 ops/s)
  5. Scale sketch (exclusive quantiles vs manual sort)
  6. Reading it
  7. Method names matter
  8. Sort once
  9. Streaming / online note
  10. Comparison to mean lab
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove