Plate 45
Counter vs dict: Tally Lab
Hands-on collections.Counter vs dict tally lab: real ops/s for Counter.update, dict get-add, and defaultdict(int) token counting on Linux localhost (lab).
Aditya Challa4 min read
Intro — what this post promises
Counting tokens: is collections.Counter faster than a manual dict or defaultdict(int)? This lab tallies N=200 000 tokens on Linux localhost across low / mid / high cardinality, comparing Counter.update, Counter(...), loop c[t]+=1, d.get, try/except, and defaultdict.
Related links:
- frozenset vs set membership localhost lab
- bytes vs bytearray localhost lab
- enum vs constants localhost lab
- itemgetter vs lambda sort localhost lab
- contextlib vs try/finally localhost lab
- itertools vs python loops localhost lab
- lru_cache hit vs miss localhost lab
- array vs list ints localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. One op = one token consumed. Affiliates: 0. Counter still wins on API (most_common, arithmetic); this is the tally hot path.
Verdict up front (mid card, vocab=5 000): Counter(tokens) ~23.4M/s, update ~22.7M, dict.get ~13.7M (update ~1.66× dict.get). Loop c[t]+=1 lags update ~2.5×. Prefer Counter(iterable) / .update, not per-item Counter math in pure Python.
Arms
| Arm | Pattern |
|---|---|
dict.get | d[t] = d.get(t, 0) + 1 |
| dict try/except | d[t]+=1 except KeyError |
defaultdict(int) | d[t] += 1 |
| Counter loop | c[t] += 1 |
Counter.update | c.update(tokens) |
Counter(tokens) | constructor from iterable |
Cardinalities: low vocab=50, mid 5 000, high 80 000.
Lab topology
Script: lab-evidence/64-counter-vs-dict-tally/results/run_lab.py.
Lead table — mid cardinality (p50)
| Arm | ops/s | ns/op |
|---|---|---|
Counter(tokens) | 23,406,456 | 42.7 |
Counter.update | 22,702,055 | 44.0 |
dict try/except | 16,058,543 | 62.3 |
defaultdict(int) | 15,503,471 | 64.5 |
dict.get | 13,676,654 | 73.1 |
Counter loop += | 9,186,252 | 108.9 |
Low vs high card (update / dict.get / ctor)
| Regime | Counter.update | dict.get | Counter() | update÷dict.get |
|---|---|---|---|---|
| low (vocab 50) | 17.6M | 19.2M | 27.0M | 0.92× |
| mid (5 000) | 22.7M | 13.7M | 23.4M | 1.66× |
| high (80 000) | 11.3M | 7.3M | 11.7M | 1.54× |
At low card, dict.get stayed competitive (update ~0.92×); ctor still led (~1.41×).
Top-k after the tally
| Arm | calls/s | µs/call |
|---|---|---|
Counter.most_common(10) | 5,165 | 193.6 |
sorted(dict.items())[:10] | 1,334 | 749.8 |
most_common ~3.9× a full sort-slice on the same mid-card map — heap-sized top-k vs sorting everything.
Reading it
Counter.update/ constructor are C-accelerated — beating a Pythonforwithc[t]+=1by ~2.5× mid-card.- defaultdict ≈ dict.get here (~1.13× mid) — pick for style, not a huge win.
- Avoid Counter as a slow dict — the loop arm is the trap.
- Use Counter when you want Counter features —
most_common, multiset ops — and feed it withupdate/ctor.
Why update beats c[t] += 1
Counter.__setitem__ from a Python loop pays interpreter overhead on every token. Counter.update and the constructor walk the iterable in C and batch the hash-table work. That is why mid-card update landed near ~23M/s while the loop arm sat near ~9M/s — same abstract algorithm, different implementation layer. If you already have a Counter and a stream, call update; if you are building from one iterable, pass it to Counter(...) directly.
Pitfalls
for t in tokens: c[t]+=1— looks idiomatic, loses the C path.- try/except for missing keys — depends on hit rate (won mid, lost low/high here).
- Sorting all items for top-10 —
most_common(k)exists. - Microbenching empty Counter creation — amortize over real N.
When to pick what
| Need | Prefer |
|---|---|
| Tally an iterable once | Counter(iterable) |
| Merge streams | c.update(...) |
| Simple counts, no Counter API | dict / defaultdict(int) |
| Top-k frequent | most_common(k) |
Reproduce
Evidence: /workspace/lab-evidence/64-counter-vs-dict-tally/results/.
Closing
Feed Counter in bulk. On this box mid-card update hit ~23M/s vs dict.get ~14M (~1.66×) and loop-+= trailed update ~2.5×. Construct or update; don’t reimplement Counter one key at a time in Python.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5; N=200k tokens. mid_card: Counter.update 22.7M vs dict.get 13.7M (~1.66x); Counter() ctor ~23.4M; defaultdict ~1.13x vs dict.get; most_common vs sorted top10 ~3.9x. Affiliates: 0. Evidence: lab-evidence/64-counter-vs-dict-tally/.
Related links
Plate 46
setdefault vs defaultdict: Insert Lab
Hands-on setdefault vs if-not-in vs defaultdict lab: real ops/s for list-append and counter default insert patterns, measured on Linux localhost (lab).
Observability & SRE · 30 Sept 2026
Plate 18
fnmatch vs re Name Filter: Localhost Lab
Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.
Observability & SRE · 30 Sept 2026
Plate 84
tarfile vs zipfile Create+Extract: Localhost Lab
Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.
Observability & SRE · 30 Sept 2026