ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 45

  1. Blog
  2. /Observability & SRE

Counter vs dict: Tally Lab

Hands-on collections.Counter vs dict tally lab: real ops/s for Counter.update, dict get-add, and defaultdict(int) token counting on Linux localhost (lab).

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — mid cardinality (p50)
  5. Low vs high card (update / dict.get / ctor)
  6. Top-k after the tally
  7. Reading it
  8. Why update beats `c[t] += 1`
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing

Intro — what this post promises

Counting tokens: is collections.Counter faster than a manual dict or defaultdict(int)? This lab tallies N=200 000 tokens on Linux localhost across low / mid / high cardinality, comparing Counter.update, Counter(...), loop c[t]+=1, d.get, try/except, and defaultdict.

Related links:

  • frozenset vs set membership localhost lab
  • bytes vs bytearray localhost lab
  • enum vs constants localhost lab
  • itemgetter vs lambda sort localhost lab
  • contextlib vs try/finally localhost lab
  • itertools vs python loops localhost lab
  • lru_cache hit vs miss localhost lab
  • array vs list ints localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. One op = one token consumed. Affiliates: 0. Counter still wins on API (most_common, arithmetic); this is the tally hot path.

Verdict up front (mid card, vocab=5 000): Counter(tokens) ~23.4M/s, update ~22.7M, dict.get ~13.7M (update ~1.66× dict.get). Loop c[t]+=1 lags update ~2.5×. Prefer Counter(iterable) / .update, not per-item Counter math in pure Python.


Arms

ArmPattern
dict.getd[t] = d.get(t, 0) + 1
dict try/exceptd[t]+=1 except KeyError
defaultdict(int)d[t] += 1
Counter loopc[t] += 1
Counter.updatec.update(tokens)
Counter(tokens)constructor from iterable

Cardinalities: low vocab=50, mid 5 000, high 80 000.


Lab topology

N = 200000 tokens/arm; repeats=9; p50 ops/s
plus 500× most_common(10) vs sorted(items)[:10]

Script: lab-evidence/64-counter-vs-dict-tally/results/run_lab.py.


Lead table — mid cardinality (p50)

Armops/sns/op
Counter(tokens)23,406,45642.7
Counter.update22,702,05544.0
dict try/except16,058,54362.3
defaultdict(int)15,503,47164.5
dict.get13,676,65473.1
Counter loop +=9,186,252108.9

Low vs high card (update / dict.get / ctor)

RegimeCounter.updatedict.getCounter()update÷dict.get
low (vocab 50)17.6M19.2M27.0M0.92×
mid (5 000)22.7M13.7M23.4M1.66×
high (80 000)11.3M7.3M11.7M1.54×

At low card, dict.get stayed competitive (update ~0.92×); ctor still led (~1.41×).


Top-k after the tally

Armcalls/sµs/call
Counter.most_common(10)5,165193.6
sorted(dict.items())[:10]1,334749.8

most_common ~3.9× a full sort-slice on the same mid-card map — heap-sized top-k vs sorting everything.


Reading it

  • Counter.update / constructor are C-accelerated — beating a Python for with c[t]+=1 by ~2.5× mid-card.
  • defaultdict ≈ dict.get here (~1.13× mid) — pick for style, not a huge win.
  • Avoid Counter as a slow dict — the loop arm is the trap.
  • Use Counter when you want Counter features — most_common, multiset ops — and feed it with update/ctor.

Why update beats c[t] += 1

Counter.__setitem__ from a Python loop pays interpreter overhead on every token. Counter.update and the constructor walk the iterable in C and batch the hash-table work. That is why mid-card update landed near ~23M/s while the loop arm sat near ~9M/s — same abstract algorithm, different implementation layer. If you already have a Counter and a stream, call update; if you are building from one iterable, pass it to Counter(...) directly.


Pitfalls

  1. for t in tokens: c[t]+=1 — looks idiomatic, loses the C path.
  2. try/except for missing keys — depends on hit rate (won mid, lost low/high here).
  3. Sorting all items for top-10 — most_common(k) exists.
  4. Microbenching empty Counter creation — amortize over real N.

When to pick what

NeedPrefer
Tally an iterable onceCounter(iterable)
Merge streamsc.update(...)
Simple counts, no Counter APIdict / defaultdict(int)
Top-k frequentmost_common(k)

Reproduce

python3 lab-evidence/64-counter-vs-dict-tally/results/run_lab.py

Evidence: /workspace/lab-evidence/64-counter-vs-dict-tally/results/.


Closing

Feed Counter in bulk. On this box mid-card update hit ~23M/s vs dict.get ~14M (~1.66×) and loop-+= trailed update ~2.5×. Construct or update; don’t reimplement Counter one key at a time in Python.

counterdefaultdictdict tallyword countcollectionspythonlocalhost labsre

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; N=200k tokens. mid_card: Counter.update 22.7M vs dict.get 13.7M (~1.66x); Counter() ctor ~23.4M; defaultdict ~1.13x vs dict.get; most_common vs sorted top10 ~3.9x. Affiliates: 0. Evidence: lab-evidence/64-counter-vs-dict-tally/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 46

    setdefault vs defaultdict: Insert Lab

    Hands-on setdefault vs if-not-in vs defaultdict lab: real ops/s for list-append and counter default insert patterns, measured on Linux localhost (lab).

    Observability & SRE · 30 Sept 2026

  • Plate 18

    fnmatch vs re Name Filter: Localhost Lab

    Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.

    Observability & SRE · 30 Sept 2026

  • Plate 84

    tarfile vs zipfile Create+Extract: Localhost Lab

    Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — mid cardinality (p50)
  5. Low vs high card (update / dict.get / ctor)
  6. Top-k after the tally
  7. Reading it
  8. Why update beats `c[t] += 1`
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove