Plate 94
statistics.quantiles vs Manual Percentile: Localhost Lab
Aditya Challa4 min read
Intro — what this post promises
Compute quartiles / percentiles on a float list via statistics.quantiles vs a sorted-index manual method. This lab reports ops/s on Linux localhost.
It is not statistics-vs-manual-mean (lab 69) — that was means; this is quantiles/percentiles.
Related links:
- statistics vs manual mean localhost lab
- graphlib topo vs manual localhost lab
- functools cache vs lru localhost lab
- tomllib vs json localhost lab
- path glob vs fnmatch localhost lab
- dataclass replace vs manual localhost lab
- zoneinfo vs utc offset localhost lab
- heapq merge vs sorted localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Uniform random floats.
Verdict up front (n=10 000): quantiles exclusive ~903.9 ops/s; manual sort+index ~904.3; quantiles on already-sorted input ~11596.9; index-only on pre-sorted ~876423.3. Sorting dominates; prefer stdlib for method semantics.
Arms
| Arm | Pattern |
|---|---|
statistics.quantiles(..., method="exclusive") | stdlib |
method="inclusive" | alternate cuts |
| quantiles on pre-sorted list | still sorts inside |
| manual sort + linear index | teaching recipe |
| index-only if pre-sorted | skip sort tax |
| nearest-rank 25/50/75 | simple percentiles |
Lab topology
Script: lab-evidence/116-statistics-quantiles-vs-manual/results/run_lab.py.
Lead table — n=10 000 (p50 ops/s)
| Arm | ops/s |
|---|---|
| manual pre-sorted index-only | 876423.3 |
| quantiles (pre-sorted input) | 11596.9 |
| manual sort + linear | 904.3 |
| statistics exclusive | 903.9 |
| nearest 25/50/75 | 807.2 |
| statistics inclusive | 759.3 |
Scale sketch (exclusive quantiles vs manual sort)
| n | statistics exclusive | manual sort+index |
|---|---|---|
| 1 000 | 10900.7 | 12343.4 |
| 10 000 | 903.9 | 904.3 |
| 100 000 | 48.9 | 64.2 |
Throughput collapses with n log n sort cost — expected.
Reading it
- Unsorted data →
statistics.quantiles(clear methods, maintained). - Already sorted stream → index math can win big (~876423.3 ops/s here) but document the rule.
- Do not confuse with mean benches (lab 69).
- Inclusive vs exclusive change cut points — pick one and stick to it in dashboards.
Method names matter
exclusive vs inclusive exist because “quartile” is not one universal formula. Shipping a blog-post index recipe into finance/report pipelines without naming the method causes silent dashboard drift. Prefer documenting statistics.quantiles(..., method=...) in the runbook.
Sort once
If you need p50, p90, p99, sort (or maintain a sorted structure) once, then index. Repeated full sorts per percentile were not the happy path on this box at n=100k (~48.9 ops/s exclusive).
Streaming / online note
These arms assume an in-memory list. True online percentiles (t-digest, P², HDR histograms) are a different design space — do not paste quantiles(list(all_samples)) into a hot telemetry path that retains every float forever. Batch windows of fixed size, then call stdlib once per window.
Comparison to mean lab
Lab 69’s mean/fsum story is O(n) without a full sort. Quartiles pay for ordering. If a dashboard only needs central tendency, do not upgrade it to quantiles “for free” — you buy n log n unless the array is already sorted for another reason.
Pitfalls
- DIY percentiles that disagree with Excel/pandas/
quantilesmethod. - Assuming pre-sorted input skips work inside
statistics.quantiles(it still sorts). - Computing many one-off percentiles each with a full sort — sort once.
- Floating ties and tiny samples (unstable cuts).
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box. Uniform floats only. Manual arms are illustrative — not bit-identical to every method=. Not numpy/pandas.
Takeaway
At 10 k floats, quantiles ~903.9 ops/s ≈ manual sort+index ~904.3. Prefer stdlib quantiles for semantics; reuse a sorted buffer if you need many cuts.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. n=10k: quantiles exclusive 903.9 ops/s; manual sort 904.3; pre-sorted index-only 876423.3. Not lab 69 mean. Affiliates: 0. Evidence: lab-evidence/116-statistics-quantiles-vs-manual/.
Related links
Plate 17
platform vs os.uname Inventory: Localhost Lab
Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 75
uuid.uuid4 vs uuid.uuid1: Localhost Lab
Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 76
cmath vs math.hypot Magnitudes: Localhost Lab
Hands-on cmath vs math.hypot magnitude ops lab: real ops/s for abs, polar, and phase, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026