Plate 76
threading.local vs tid-dict: Localhost Lab
Hands-on threading.local vs tid-keyed dict vs module-global lab: real ops/s for per-thread state access patterns, measured on Linux localhost for SREs.
Aditya Challa4 min read
Intro — what this post promises
Where should per-thread state live — threading.local, a dict keyed by threading.get_ident(), or a module global? This lab measures ops/s for a tight increment loop on Linux localhost, single-thread and with a small thread pool.
Related links:
- threading event vs condition localhost lab
- threadpoolexecutor vs sequential localhost lab
- asyncio gather vs taskgroup localhost lab
- processpoolexecutor vs sequential localhost lab
- setdefault vs defaultdict localhost lab
- csv reader vs split localhost lab
- pickle vs json roundtrip localhost lab
- mmap vs read scan localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. No Docker. Selectors/epoll already covered in epoll vs select FD_SETSIZE localhost lab — this post is TLS/state access only.
Verdict up front (2 M increments, single thread): global dict ~22.11 Mops/s; tid-dict ~15.16 Mops/s (~1.27× vs local); threading.local ~11.93 Mops/s. Under 8 workers, tid-dict ~12.49 Mops/s vs TLS ~9.62 vs locked global ~6.52 (correctness tax).
Arms
| Arm | Pattern |
|---|---|
threading.local | tls.counter += 1 (per-thread attrs) |
| tid-dict | d[get_ident()] += 1 (own key only) |
| module global dict | shared GLOBAL_STATE["counter"] += 1 |
| global + Lock (MT) | same counter behind threading.Lock |
Lab topology
Script: lab-evidence/84-threading-local-vs-dict/results/run_lab.py.
Lead table — single thread (p50)
| Arm | Mops/s | vs threading.local |
|---|---|---|
| module global dict | 22.11 | 1.85× |
| tid-dict get/set | 15.16 | 1.27× |
| threading.local | 11.93 | 1.00× |
threading.local is convenient — not free. A plain dict keyed by thread id won on this microbench.
Multi-thread (total ops/s)
| Workers | TLS Mops/s | tid-dict Mops/s | locked global Mops/s |
|---|---|---|---|
| 4 | 9.31 | 11.78 | 6.63 |
| 8 | 9.62 | 12.49 | 6.52 |
Locked shared global is the slowest here — expected when every increment contends. tid-dict with per-thread keys needs no lock for this pattern and stayed ahead of TLS.
Reading it
- Globals are fastest and wrong for per-thread semantics (shared mutable state / races without a lock).
threading.localis the readable API for request/context bags; expect a measurable tax vs a hand-rolled tid map on tiny ops.- tid-dict is fine when you control lifecycle (delete keys on thread exit) and only touch your own key.
- Lock the global if the state is truly shared — throughput collapses vs TLS/tid maps.
Correctness vs speed
This lab’s increment is the hottest possible access. In real apps, TLS lookup is rarely the bottleneck next to I/O. Prefer threading.local (or contextvars for async) for clarity unless profiles show TLS in the top frames.
Why not selectors for lab 84
epoll vs select FD_SETSIZE already measured readiness wait scaling. Re-baking poll/epoll here would cannibalize that post — TLS access is the fresher gap.
contextvars footnote
Async code should prefer contextvars over threading.local — TLS does not follow asyncio tasks. This lab stays on threads; if your stack is async, measure contextvar get/set separately before copying these Mops/s.
Pitfalls
- Creating a new
threading.local()inside each task instead of one shared TLS object (defeats the point). - Using a shared dict without per-thread keys or a lock (races).
- Leaking tid-dict entries when threads die and restart with new idents.
- Treating module globals as “thread-safe” because CPython’s GIL exists — compound
+=on containers is still a logic race for app state.
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box. Microbench increments only. Not contextvars, not free-threaded Python builds. Numbers move with GIL changes and CPU count.
Takeaway
For per-thread counters, tid-dict (~15.16 Mops/s) beat threading.local (~11.93 Mops/s) on this host; globals (~22.11 Mops/s) are faster still but are shared state. Under threads, locked globals (~6.52 Mops/s ) pay the correctness tax — use TLS or own-key maps for thread-private data.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. Single-thread 2M incs: global 22.11 Mops/s; tid-dict 15.16; threading.local 11.93. MT w8: tid-dict 12.49; TLS 9.62; locked global 6.52. Affiliates: 0. Evidence: lab-evidence/84-threading-local-vs-dict/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 71
html.escape vs Manual Replace: Localhost Lab
A hands-on localhost lab comparing html.escape with chained str.replace for safe HTML escaping.
Observability & SRE · 30 Sept 2026
Plate 17
difflib vs set Ops Similarity: Localhost Lab
Hands-on difflib.SequenceMatcher vs set Jaccard token similarity: real ops/s on token lists, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 30 Sept 2026