ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 46

  1. Blog
  2. /Observability & SRE

setdefault vs defaultdict: Insert Lab

Hands-on setdefault vs if-not-in vs defaultdict lab: real ops/s for list-append and counter default insert patterns, measured on Linux localhost (lab).

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — list-append groupby (p50)
  5. Scalar counters (p50)
  6. Reading it
  7. setdefault vs defaultdict in APIs
  8. Pitfalls
  9. When to pick what
  10. Mixed-hit takeaway
  11. Reproduce
  12. Closing

Intro — what this post promises

Insert-or-default on a dict: is setdefault, an if k not in d, or collections.defaultdict fastest? This lab times list-append groupby and scalar counter patterns on Linux localhost under miss-heavy / mixed / hit-heavy key mixes.

Related links:

  • Counter vs dict tally localhost lab
  • functools partial vs lambda localhost lab
  • zip vs index pairing localhost lab
  • frozenset vs set membership localhost lab
  • enum vs constants localhost lab
  • contextlib vs try/finally localhost lab
  • statistics vs manual mean localhost lab
  • bisect vs linear lookup localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. N=200,000 key ops/arm. List arms seed 5 000 existing keys. Affiliates: 0. Complements Counter tally with the “ensure bucket then mutate” pattern.

Verdict up front (mixed list-append): defaultdict ~5.11M/s ≈ setdefault ~5.08M (~1.01×); hit-heavy defaultdict ~1.08× setdefault. Scalar hit: d.get/defaultdict(int) beat setdefault++= by ~1.36× / ~1.32×. Pick defaultdict for ongoing groupby; use d[k]=d.get(k,0)+1 for counters.


Arms

ArmPattern
setdefaultd.setdefault(k, []).append(1)
if-not-inif k not in d: d[k]=[] then append
try/exceptappend, catch KeyError
defaultdict(list)d[k].append(1)
get-orlst=d.get(k); ...
scalarsetdefault/get/defaultdict(int) counters

Hit rates: miss 10%, mixed 50%, hit 90%.


Lab topology

N = 200000 ops; list seed = 5000 keys
scalar miss = all-new keys; scalar hit = 1000 hot keys

Script: lab-evidence/72-setdefault-vs-defaultdict/results/run_lab.py.


Lead table — list-append groupby (p50)

Armmiss-heavymixedhit-heavy
setdefault3.17M5.08M13.58M
if-not-in2.82M4.81M11.84M
defaultdict2.81M5.11M14.73M
get-or2.96M5.16M13.55M
try/except2.30M4.17M13.21M

Scalar counters (p50)

Armall-misshit-heavy
get-add7.90M20.50M
setdefault then +=6.78M15.11M
if-not-in6.40M15.65M
defaultdict(int)5.88M19.90M

Note: setdefault + += does an extra lookup after insert — d[k]=d.get(k,0)+1 avoids that.


Reading it

  • List groupby — setdefault ≈ defaultdict on mixed; defaultdict pulls ahead when hits dominate (~1.08×).
  • Miss-heavy list — setdefault slightly ahead of defaultdict here (dd÷sd ~0.89×); gaps are small.
  • Counters — prefer get/defaultdict(int), not setdefault+mutate for ints.
  • try/except — loses on miss-heavy (exception path); fine when hits are near-certain.

setdefault vs defaultdict in APIs

defaultdict changes missing-key behavior for the lifetime of the object — fine inside a function, surprising if you return it to callers who expect KeyError. setdefault keeps a plain dict and only inserts when you ask. For library boundaries, prefer setdefault (or build with defaultdict then dict(d)). For tight internal grouping loops, defaultdict’s __getitem__ path is the ergonomic default and often the speed default on hit-heavy mixes.


Pitfalls

  1. setdefault for int counters — double lookup with +=.
  2. Leaking defaultdict into APIs that expect plain dict — convert or document.
  3. try/except as control flow on cold keys — expensive misses.
  4. Microbenching without hit-rate context — rankings flip with mix.

When to pick what

NeedPrefer
Ongoing list/set bucketsdefaultdict(list) / setdefault
Int/float countersd[k]=d.get(k,0)+1 or defaultdict(int) / Counter
One-shot ensuresetdefault
Plain dict requiredsetdefault / if-not-in

Mixed-hit takeaway

Across miss / mixed / hit, no single helper wins every cell by a wide margin on list buckets — the story is “same ballpark, pick clarity,” except try/except on cold keys and setdefault++= for ints. Profile the hit rate you actually see in production logs before rewriting idioms for a 5% microbench delta.


Reproduce

python3 lab-evidence/72-setdefault-vs-defaultdict/results/run_lab.py

Evidence: /workspace/lab-evidence/72-setdefault-vs-defaultdict/results/.


Closing

defaultdict for buckets; get-add for counters. On this box mixed list-append setdefault and defaultdict tied near ~5.1M/s, hit-heavy favored defaultdict by ~1.08×, and scalar get-add led setdefault by ~1.36×. Match the helper to whether the default is a container or a number.

setdefaultdefaultdictdictif not inpythonlocalhost labsreinsert-or-default

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; N=200000. Mixed list-append: defaultdict 5.11M ≈ setdefault 5.08M (~1.01x); hit-heavy defaultdict ~1.08x. Scalar hit: get-add/defaultdict beat setdefault (~1.36x/~1.32x). Affiliates: 0. Evidence: lab-evidence/72-setdefault-vs-defaultdict/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 45

    Counter vs dict: Tally Lab

    Hands-on collections.Counter vs dict tally lab: real ops/s for Counter.update, dict get-add, and defaultdict(int) token counting on Linux localhost (lab).

    Observability & SRE · 30 Sept 2026

  • Plate 18

    fnmatch vs re Name Filter: Localhost Lab

    Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.

    Observability & SRE · 30 Sept 2026

  • Plate 84

    tarfile vs zipfile Create+Extract: Localhost Lab

    Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — list-append groupby (p50)
  5. Scalar counters (p50)
  6. Reading it
  7. setdefault vs defaultdict in APIs
  8. Pitfalls
  9. When to pick what
  10. Mixed-hit takeaway
  11. Reproduce
  12. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove