Plate 85
math.fsum vs sum Float Totals: Localhost Lab
Hands-on math.fsum vs builtin sum for float totals: real items/s plus cancellation accuracy, measured on Linux localhost in this hands-on lab for SREs.
Aditya Challa4 min read
Intro — what this post promises
Total floats with math.fsum vs builtin sum vs a naive += loop. This lab reports items/s on Linux localhost, plus a cancellation accuracy case that still bites on Python 3.13.
Related links:
- fractions vs float localhost lab
- statistics quantiles vs manual localhost lab
- decimal vs float sum localhost lab
- random choices vs sample localhost lab
- graphlib topo vs manual localhost lab
- functools cache vs lru localhost lab
- tomllib vs json localhost lab
- itertools batched vs chunk localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Not Decimal (lab 85) and not Fraction (lab 122) — this is IEEE float summation only.
Verdict up front (n=100 000 uniform): sum ~205.04 Mitems/s; math.fsum ~188.64; naive loop ~84.48. Mixed magnitudes: sum ~206.94 vs fsum ~46.51. Accuracy: sum([1, 1e100, 1, -1e100]) → 0.0; math.fsum(...) → 2.0.
Arms
| Arm | Pattern |
|---|---|
sum(data) | builtin float sum |
math.fsum(data) | precise float sum |
naive s += x | pure Python loop |
| mixed-magnitude variants | stress partials |
statistics.fmean * n | related mean path |
Seven rounds, p50 items/s. Uniform draws keep magnitudes similar; mixed-exponent draws force larger intermediate cancellations.
Lab topology
Script: lab-evidence/123-math-fsum-vs-sum/results/run_lab.py.
Lead table — n=100 000 (p50 Mitems/s)
| Arm | Mitems/s |
|---|---|
| sum builtin | 205.04 |
| math.fsum | 188.64 |
| statistics.fmean×n | 187.1 |
| naive loop | 84.48 |
| sum mixed mag | 206.94 |
| fsum mixed mag | 46.51 |
On uniform data, sum and fsum are close; the naive loop loses badly to interpreter overhead.
Scale (uniform)
| n | sum | fsum |
|---|---|---|
| 10 000 | 209.33 | 194.45 |
| 100 000 | 205.04 | 188.64 |
| 1 000 000 | 149.79 | 105.49 |
Throughput stays in the same ballpark as n grows — both paths are C-backed; the story is accuracy under cancellation, not asymptotic surprise.
Cancellation footgun
For [1, 1e100, 1, -1e100]: sum → 0.0, fsum → 2.0. Prefer math.fsum when large partials can cancel. Keep the classic mixed int/float literals in regression tests — all-float versions can take a different sum path on 3.13 and hide the bug.
Mixed-magnitude cost
On mixed exponents, fsum slowed to ~46.51 Mitems/s vs sum ~206.94 — precision work is not free. That is the trade: pay for shepherds of partials when correctness matters more than peak items/s.
Reading it for SRE work
- Hot path, similar magnitudes →
sum. - Large dynamic range / cancellation risk (log totals, compensated sensors) →
math.fsum. - Never hand-roll
+=for speed — it lost to both C paths. - Exact decimals still want Decimal (lab 85) or Fraction (lab 122).
Document which reducer you pick in runbooks so on-call does not “optimize” away a deliberate fsum later.
Pitfalls
- Assuming
sumandfsumalways agree. - Using fsum then wiping precision in display formats.
- Confusing with lab 85 Decimal arithmetic.
- Testing only all-float literals for cancellation demos on 3.13.
- Replacing
fsumwithsumin a hot path without an accuracy fixture.
Reproduce
Evidence: summary.json, summary.txt under lab-evidence/123-math-fsum-vs-sum/results/.
Limits
One Linux box. CPython float summation only. Not GPU reductions, not multiprecision.
Takeaway
At 100 k uniforms, sum ~205.04 Mitems/s edged fsum ~188.64. Use math.fsum when cancellation matters — classic case: sum 0.0 vs fsum 2.0.
Lab evidence
What I found running this
Ran the package lab on Linux localhost with CPython 3.13.5 on 1 Oct 2026 IST. Benchmarked seven p50 rounds for uniform and mixed-magnitude floats at 10k, 100k, and 1M items. Verified cancellation: sum returned 0.0 while math.fsum returned 2.0; affiliates 0.
Related links
Plate 77
fractions.Fraction vs float: Localhost Lab
Hands-on fractions.Fraction vs float for exact ratios: real ops/s and exactness checks, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 57
memoryview vs bytes Slice: Localhost Lab
1 Oct 2026
Plate 82
shlex.split vs str.split: Localhost Lab
1 Oct 2026