Plate 46
setdefault vs defaultdict: Insert Lab
Hands-on setdefault vs if-not-in vs defaultdict lab: real ops/s for list-append and counter default insert patterns, measured on Linux localhost (lab).
Aditya Challa4 min read
Intro — what this post promises
Insert-or-default on a dict: is setdefault, an if k not in d, or collections.defaultdict fastest? This lab times list-append groupby and scalar counter patterns on Linux localhost under miss-heavy / mixed / hit-heavy key mixes.
Related links:
- Counter vs dict tally localhost lab
- functools partial vs lambda localhost lab
- zip vs index pairing localhost lab
- frozenset vs set membership localhost lab
- enum vs constants localhost lab
- contextlib vs try/finally localhost lab
- statistics vs manual mean localhost lab
- bisect vs linear lookup localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. N=200,000 key ops/arm. List arms seed 5 000 existing keys. Affiliates: 0. Complements Counter tally with the “ensure bucket then mutate” pattern.
Verdict up front (mixed list-append): defaultdict ~5.11M/s ≈ setdefault ~5.08M (~1.01×); hit-heavy defaultdict ~1.08× setdefault. Scalar hit: d.get/defaultdict(int) beat setdefault++= by ~1.36× / ~1.32×. Pick defaultdict for ongoing groupby; use d[k]=d.get(k,0)+1 for counters.
Arms
| Arm | Pattern |
|---|---|
| setdefault | d.setdefault(k, []).append(1) |
| if-not-in | if k not in d: d[k]=[] then append |
| try/except | append, catch KeyError |
| defaultdict(list) | d[k].append(1) |
| get-or | lst=d.get(k); ... |
| scalar | setdefault/get/defaultdict(int) counters |
Hit rates: miss 10%, mixed 50%, hit 90%.
Lab topology
Script: lab-evidence/72-setdefault-vs-defaultdict/results/run_lab.py.
Lead table — list-append groupby (p50)
| Arm | miss-heavy | mixed | hit-heavy |
|---|---|---|---|
| setdefault | 3.17M | 5.08M | 13.58M |
| if-not-in | 2.82M | 4.81M | 11.84M |
| defaultdict | 2.81M | 5.11M | 14.73M |
| get-or | 2.96M | 5.16M | 13.55M |
| try/except | 2.30M | 4.17M | 13.21M |
Scalar counters (p50)
| Arm | all-miss | hit-heavy |
|---|---|---|
| get-add | 7.90M | 20.50M |
| setdefault then += | 6.78M | 15.11M |
| if-not-in | 6.40M | 15.65M |
| defaultdict(int) | 5.88M | 19.90M |
Note: setdefault + += does an extra lookup after insert — d[k]=d.get(k,0)+1 avoids that.
Reading it
- List groupby — setdefault ≈ defaultdict on mixed; defaultdict pulls ahead when hits dominate (~1.08×).
- Miss-heavy list — setdefault slightly ahead of defaultdict here (dd÷sd ~0.89×); gaps are small.
- Counters — prefer
get/defaultdict(int), not setdefault+mutate for ints. - try/except — loses on miss-heavy (exception path); fine when hits are near-certain.
setdefault vs defaultdict in APIs
defaultdict changes missing-key behavior for the lifetime of the object — fine inside a function, surprising if you return it to callers who expect KeyError. setdefault keeps a plain dict and only inserts when you ask. For library boundaries, prefer setdefault (or build with defaultdict then dict(d)). For tight internal grouping loops, defaultdict’s __getitem__ path is the ergonomic default and often the speed default on hit-heavy mixes.
Pitfalls
setdefaultfor int counters — double lookup with+=.- Leaking defaultdict into APIs that expect plain
dict— convert or document. - try/except as control flow on cold keys — expensive misses.
- Microbenching without hit-rate context — rankings flip with mix.
When to pick what
| Need | Prefer |
|---|---|
| Ongoing list/set buckets | defaultdict(list) / setdefault |
| Int/float counters | d[k]=d.get(k,0)+1 or defaultdict(int) / Counter |
| One-shot ensure | setdefault |
| Plain dict required | setdefault / if-not-in |
Mixed-hit takeaway
Across miss / mixed / hit, no single helper wins every cell by a wide margin on list buckets — the story is “same ballpark, pick clarity,” except try/except on cold keys and setdefault++= for ints. Profile the hit rate you actually see in production logs before rewriting idioms for a 5% microbench delta.
Reproduce
Evidence: /workspace/lab-evidence/72-setdefault-vs-defaultdict/results/.
Closing
defaultdict for buckets; get-add for counters. On this box mixed list-append setdefault and defaultdict tied near ~5.1M/s, hit-heavy favored defaultdict by ~1.08×, and scalar get-add led setdefault by ~1.36×. Match the helper to whether the default is a container or a number.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5; N=200000. Mixed list-append: defaultdict 5.11M ≈ setdefault 5.08M (~1.01x); hit-heavy defaultdict ~1.08x. Scalar hit: get-add/defaultdict beat setdefault (~1.36x/~1.32x). Affiliates: 0. Evidence: lab-evidence/72-setdefault-vs-defaultdict/.
Related links
Plate 45
Counter vs dict: Tally Lab
Hands-on collections.Counter vs dict tally lab: real ops/s for Counter.update, dict get-add, and defaultdict(int) token counting on Linux localhost (lab).
Observability & SRE · 30 Sept 2026
Plate 18
fnmatch vs re Name Filter: Localhost Lab
Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.
Observability & SRE · 30 Sept 2026
Plate 84
tarfile vs zipfile Create+Extract: Localhost Lab
Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.
Observability & SRE · 30 Sept 2026