Plate 99
flock Contention Lab: Exclusive, Shared, and LOCK_NB Fail Rates
Hands-on flock lab: uncontended LOCK_EX p50 0.33 us; waiter ~50 ms hold; w8 p50 ~36 ms; LOCK_NB 99% fail; LOCK_SH ~6.7x wall. Real numbers, no Docker.
Aditya Challa5 min read
On this page
- Intro — what this post promises
- What we compared
- Lab topology
- Arm A — uncontended LOCK\_EX
- Arm B — holder blocks waiter (hold = 50 ms)
- Arm C — contended LOCK\_EX (hold = 5 ms, 25 rounds/worker)
- Arm D — LOCK\_NB fail rate
- Arm E — LOCK\_SH overlap vs LOCK\_EX serialize
- When exclusive still wins the design
- Pitfalls we hit (or avoided)
- Practical checklist
- Verdict
Intro — what this post promises
Two processes, one lockfile, and a queue you cannot see in top. flock(LOCK_EX) is the classic Linux advisory exclusive lock. Uncontended it is nearly free. Contended it becomes a serialized wait queue. Shared locks (LOCK_SH) are supposed to overlap. LOCK_NB is supposed to fail fast instead of sleeping.
This is a hands-on lab with measured numbers:
- Uncontended
LOCK_EXacquire latency (microseconds). - One holder, one waiter — block time ≈ hold time.
- Contended exclusive acquire p50 as worker count grows.
LOCK_NBfail rate under contention.LOCK_SHwall-time overlap vs exclusive serialize.
Related links:
- Pipe vs tmpfile IPC localhost lab
- ulimit soft vs hard file descriptors lab
- mmap vs read (+ O_DIRECT) localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 fcntl.flock on a local file (run/lab.lock). ProcessPool waiters. No NFS. No Docker. No GPU. No API keys. Affiliates: 0.
Verdict up front: uncontended p50 ≈ 0.33 µs; waiter behind a 50 ms hold ≈ 50.4 ms; contended w=8 p50 ≈ 36 ms (5 ms holds); LOCK_NB w=8 fail ≈ 99%; LOCK_SH wall ~6.7× vs serial hold.
What we compared
| Arm | API | Metric |
|---|---|---|
| A | LOCK_EX alone | acquire µs |
| B | holder + waiter | waiter block ms ≈ hold |
| C | N processes LOCK_EX | acquire p50 under queueing |
| D | LOCK_EX|LOCK_NB | fail % |
| E | LOCK_SH vs LOCK_EX | wall overlap vs serialize |
Related links:
Lab topology
Arm A — uncontended LOCK_EX
| Metric | Value |
|---|---|
| p50 | 0.33 µs |
| p95 | 0.35 µs |
| mean | 0.35 µs |
Without a competitor, exclusive flock is noise-level cheap on a local filesystem.
Arm B — holder blocks waiter (hold = 50 ms)
| Metric | Waiter acquire |
|---|---|
| p50 | 50.4 ms |
| p95 | 52.6 ms |
| mean | 50.8 ms |
The waiter slept for essentially the holder’s critical section. Contended flock is a queue, not a spin that burns CPU in our process (the waiter is blocked in the kernel).
Arm C — contended LOCK_EX (hold = 5 ms, 25 rounds/worker)
| Workers | Acquire p50 ms | p95 ms | mean ms |
|---|---|---|---|
| 2 | 4.97 | 7.13 | 4.90 |
| 4 | 15.2 | 20.6 | 15.3 |
| 8 | 35.9 | 44.7 | 35.1 |
More waiters → longer average queue. Rough intuition: with hold H and W workers, waits climb toward ~(W−1)·H under steady contention (8 workers × 5 ms → tens of ms — matches the table).
Arm D — LOCK_NB fail rate
| Workers | Got | Fail | Fail % |
|---|---|---|---|
| 1 | 60 | 0 | 0% |
| 8 | 4 | 476 | 99.2% |
Non-blocking exclusive locks almost never succeed once eight processes hammer an 8 ms critical section. Use LOCK_NB when fail-fast / try-later is the product behavior — not when you need eventual entry.
Arm E — LOCK_SH overlap vs LOCK_EX serialize
8 processes × 15 rounds × 10 ms hold (ideal serial hold = 1.20 s):
| Mode | Acquire p50 | Wall s | vs ideal serial |
|---|---|---|---|
| LOCK_SH | 1.92 µs | 0.178 | ~6.7× faster wall |
| LOCK_EX | 71.0 ms | 1.24 | ~1.03× ideal |
Shared locks overlapped almost to core count. Exclusive locks serialized to the hold budget. That is the whole design choice in one table.
Related links:
When exclusive still wins the design
- Mutating a shared file / PID file / deploy stamp — readers must not see torn state.
- Single-flight cron / job guards — one winner, others exit (
LOCK_NB) or wait. - Not a distributed lock — this lab is local
flock, not Redis/etcd/NFS edge cases.
Use LOCK_SH when many readers can safely overlap; measure wall time — Arm E is the template.
Pitfalls we hit (or avoided)
- Treating uncontended µs as the production number — Arm C is the one that hurts.
- Expecting LOCK_NB to “mostly work” under load — 99% fail at w=8.
- Assuming flock on NFS matches local — we did not test NFS; do not extrapolate.
- Forgetting advisory means cooperative — a process that never locks is unconstrained.
- Calling this a mutex benchmark for threads in one process — these are process waiters on a file.
Practical checklist
- Critical section duration × waiter count ≈ user-visible stall — budget hold time first.
- Fail-fast paths:
LOCK_NB+ retry/jitter; do not spin forever in userspace. - Reader-heavy: prefer
LOCK_SH(or a different shared-state design); prove with wall overlap. - Report hold ms + worker count + p50 wait + NB fail% with every flock claim.
- Keep lockfiles on local disk for this semantic; validate separately before trusting NFS.
Verdict
Local fcntl.flock: uncontended LOCK_EX p50 ≈ 0.33 µs; a 50 ms holder blocked the waiter ≈ 50.4 ms; under 5 ms holds, w=8 acquire p50 ≈ 36 ms; LOCK_NB failed ~99% at w=8; LOCK_SH delivered ~6.7× wall speedup vs serial while LOCK_EX matched ideal serialize (~1.03×).
Evidence path on the lab box: lab-evidence/27-flock-contention/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5 fcntl.flock on local file. Uncontended LOCK_EX p50 0.33 µs. Holder-block waiter hold=50 ms → waiter p50 50.4 ms. Contended LOCK_EX hold=5 ms: w2 p50 5.0 ms; w4 15.2 ms; w8 35.9 ms. LOCK_NB: w1 fail 0%; w8 fail 99.2%. LOCK_SH 8×15×10 ms: wall 0.178 s (~6.7× vs 1.2 s serial); LOCK_EX wall 1.24 s ≈ ideal. No Docker. Affiliates: 0. Evidence: lab-evidence/27-flock-contention/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026