ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 58

  1. Blog
  2. /Observability & SRE

groupby vs Manual Group: Localhost Lab

Hands-on itertools.groupby vs manual dict-of-lists: real records/s grouping pre-sorted key runs, measured on Linux localhost today in this lab for SREs.

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 50 000 records (p50 Mrec/s)
  5. Scale sketch (groupby → dict vs defaultdict)
  6. Unsorted footgun
  7. Sort tax reminder
  8. Reading it
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

Group a pre-sorted list of (key, value) records: itertools.groupby vs a manual run-length scan vs defaultdict(list). This lab reports records/s on Linux localhost, and shows what happens if you feed unsorted data to groupby.

It is not Counter tallying (lab 64) or chain/flatten (lab 67), and not the itertools micro-loop bake-off (lab 55).

Related links:

  • counter vs dict tally localhost lab
  • itertools chain vs flatten localhost lab
  • itertools vs python loops localhost lab
  • setdefault vs defaultdict localhost lab
  • fnmatch vs re localhost lab
  • futures as completed vs wait localhost lab
  • hmac compare digest localhost lab
  • argparse vs sys argv localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. No Docker. Input sorted unless noted.

Verdict up front (50 000 records / 200 keys): groupby consume ~24.21 Mrec/s; manual sorted runs ~18.09; defaultdict ~19.08; groupby→dict lists ~17.24. On shuffled input, naive d[k]=list(grp) lost 49800 / 50000 values.


Arms

ArmPattern
groupby consumeiterate groups, discard values
groupby → dict of listsmaterialize each group
manual sorted runscontiguous key scan
defaultdict(list)works unsorted
dict.setdefaultsame idea, slower API

Lab topology

n×keys: 10k/100, 50k/200, 200k/500 · sorted · 7 rounds · p50
unsorted demo: shuffle 50k, compare naive groupby vs defaultdict
metric: records/s = n / p50_seconds

Script: lab-evidence/98-groupby-vs-manual/results/run_lab.py.


Lead table — 50 000 records (p50 Mrec/s)

ArmMrec/s
groupby consume24.21
manual sorted runs18.09
defaultdict(list)19.08
groupby → dict lists17.24
setdefault lists15.19

At this scale, consume is fastest because it never allocates value lists. Once you materialize, manual runs, groupby→dict, and defaultdict sit in a tight band — list appends dominate.


Scale sketch (groupby → dict vs defaultdict)

Configgroupby→dictdefaultdictgroupby consume
10k / 100 keys27.4125.0543.23
50k / 20017.2419.0824.21
200k / 5006.877.918.71

Throughput falls as group and list sizes grow — expected allocator pressure, not a mystery.


Unsorted footgun

Naive assign-per-run on shuffled data kept only 200 values (lost 49800). Merging runs or using defaultdict kept all 50000. groupby requires sorted (or already-grouped) input.


Sort tax reminder

If your pipeline is unsorted, add sorted(..., key=) cost before celebrating groupby. For one-shot grouping of random keys, defaultdict often wins end-to-end even when per-record rates look close on a pre-sorted fixture. Manual run-length scans assume contiguity the same way groupby does.


Reading it

  • Prefer groupby when data is already sorted and you stream group iterators.
  • defaultdict(list) wins when order is arbitrary or you do not want a sort tax.
  • Materializing lists dominates — consume iterators when you can.
  • Manual run-length scan ≈ groupby→dict when both build lists.
  • setdefault trails defaultdict on this box for the same list-building work.

Pitfalls

  • Using groupby on unsorted data (silent wrong merges / lost values).
  • Assuming groupby output keys are unique without sorting.
  • Sorting a huge list just to use groupby when defaultdict would do.
  • Confusing grouping with Counter (counts only — lab 64).
  • Measuring only consume throughput then shipping code that builds full dicts.

Reproduce

python3 lab-evidence/98-groupby-vs-manual/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box. Synthetic int values. Sort cost not included in sorted-arm timings (fixture pre-sorted). Warmup discarded; p50 of seven rounds.


Takeaway

On sorted 50 k records, groupby consume ~24.21 Mrec/s led; building dicts sat ~17.24–19.08 Mrec/s. Sort first or use defaultdict — unsorted naive groupby dropped 49800 values here.

itertools.groupbygroupby sorteddefaultdict listgroup recordspython itertoolslocalhost labsrerecords/s

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. 50k/200keys: groupby_consume 24.21 Mrec/s; manual_runs 18.09; defaultdict 19.08; groupby_dict 17.24. Unsorted naive lost 49800/50000. Affiliates: 0. Evidence: lab-evidence/98-groupby-vs-manual/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 50 000 records (p50 Mrec/s)
  5. Scale sketch (groupby → dict vs defaultdict)
  6. Unsorted footgun
  7. Sort tax reminder
  8. Reading it
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove