Plate 80
Event vs Condition vs Barrier: Localhost Lab
Aditya Challa4 min read
Intro — what this post promises
Need to wake N worker threads from one controller — threading.Event, Condition.notify_all / notify(1), or a Barrier? This lab measures wake latency and wakes/s for broadcast-style sync on Linux localhost.
Related links:
- threadpoolexecutor vs sequential localhost lab
- processpoolexecutor vs sequential localhost lab
- asyncio gather vs taskgroup localhost lab
- process vs thread pool GIL localhost lab
- mmap vs read scan localhost lab
- deque vs list queue localhost lab
- csv reader vs split localhost lab
- pickle vs json roundtrip localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Pure stdlib threading. Affiliates: 0. No Docker. Timing starts after all waiters are known to be waiting (no lost-wakeup races in the Condition arms).
Verdict up front (128 waiters broadcast): Event ~8.071 ms (~15860 wakes/s); Condition.notify_all ~7.068 ms (~18110/s); Barrier ~9.247 ms. Event ping-pong handoff ~16325 ns (~61254/s).
Arms
| Arm | Pattern |
|---|---|
| Event broadcast | N waiters on Event.wait; one set() |
Condition notify_all | N waiters in cv.wait; one notify_all |
Condition notify(1) × N | wake one-by-one after all parked |
| Barrier | N workers + main Barrier.wait rendezvous |
| Event ping-pong | two threads alternate Events (handoff latency) |
Lab topology
Script: lab-evidence/83-threading-event-vs-condition/results/run_lab.py.
Lead table — broadcast wake (p50)
| N | Event ms | Event wakes/s | notify_all ms | notify_all /s | Barrier ms |
|---|---|---|---|---|---|
| 4 | 0.18 | 22224 | 0.097 | 41136 | 0.136 |
| 16 | 0.727 | 22016 | 0.567 | 28230 | 0.815 |
| 64 | 4.076 | 15702 | 3.548 | 18040 | 4.874 |
| 128 | 8.071 | 15860 | 7.068 | 18110 | 9.247 |
notify(1) × N (after all parked) at N=128: ~1.194 ms — useful when you must release workers one at a time, not as a faster broadcast.
Ping-pong latency
Two threads alternating Event signals: ~16325 ns/handoff (~61254 handoffs/s). That is the cost of a round-trip wake for a simple flag protocol — fine for control planes, not a substitute for a lock-free queue in a hot path.
Reading it
- Event is the simplest broadcast: sticky
set()is hard to misuse for one-shot “go” flags. - Condition.notify_all matched or slightly beat Event at large N here (~18110 vs ~15860 wakes/s at 128) and wins when you need a predicate loop (
while not ready: cv.wait()). - Barrier is for rendezvous (everyone meets), not “one signals many” — wall time includes all parties arriving.
notify(1)× N is a different pattern (fair one-by-one release), not a drop-in fasternotify_all.
When to pick which
| Need | Prefer |
|---|---|
| One-shot “start now” flag | Event |
| Wait until predicate / shared state | Condition |
| All threads reach a phase gate | Barrier |
| Hand off work one worker at a time | notify(1) (or a Queue) |
Lost-wakeup note
Bare Condition.notify before the waiter enters wait() loses the signal. This lab parks everyone (counter + cv.wait handshake) before timing notifies. Production code should use predicate loops, not sleep-and-hope.
Scaling sketch
From N=4 → 128, Event wake wall grew ~45× while wakes/s stayed in a ~15–22k band once N≥64. That is thread create/join and scheduler fan-out, not a futex-only number. For hotter fan-out, prefer a work queue over waking idle thread armies.
Pitfalls
- Using Barrier when you meant broadcast (or vice versa).
- Forgetting to
clear()an Event between phases. - Calling
notify(1)without holding the lock / without predicates. - Measuring wake time without ensuring waiters are blocked first (optimistic numbers).
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box, process-local threads only. Not multiprocessing Events. Not futex microbenchmarks. Scheduler noise dominates at sub-ms wakes.
Takeaway
For waking N waiters, Event and Condition.notify_all are in the same band (~8.071–7.068 ms at N=128, ~15–18k wakes/s). Use Event for simple flags, Condition for predicates, Barrier for rendezvous, and treat ping-pong ~16325 ns as the two-thread handoff floor — not a throughput engine.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Ran Python 3.13.5 on Linux localhost with pure stdlib threading. At N=128 broadcast waiters, Event took 8.071 ms (~15860 wakes/s), Condition.notify_all took 7.068 ms (~18110/s), and Barrier took 9.247 ms. Event ping-pong measured about 16325 ns per handoff (~61254/s). Waiters were confirmed parked before timing; no Docker, multiprocessing, or affiliate links were used.
Related links
Plate 76
threading.local vs tid-dict: Localhost Lab
Hands-on threading.local vs tid-keyed dict vs module-global lab: real ops/s for per-thread state access patterns, measured on Linux localhost for SREs.
Observability & SRE · 30 Sept 2026
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026