ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 99

  1. Blog
  2. /Observability & SRE

flock Contention Lab: Exclusive, Shared, and LOCK_NB Fail Rates

Hands-on flock lab: uncontended LOCK_EX p50 0.33 us; waiter ~50 ms hold; w8 p50 ~36 ms; LOCK_NB 99% fail; LOCK_SH ~6.7x wall. Real numbers, no Docker.

Aditya Challa·30 September 2026·5 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — uncontended LOCK\_EX
  5. Arm B — holder blocks waiter (hold = 50 ms)
  6. Arm C — contended LOCK\_EX (hold = 5 ms, 25 rounds/worker)
  7. Arm D — LOCK\_NB fail rate
  8. Arm E — LOCK\_SH overlap vs LOCK\_EX serialize
  9. When exclusive still wins the design
  10. Pitfalls we hit (or avoided)
  11. Practical checklist
  12. Verdict

Intro — what this post promises

Two processes, one lockfile, and a queue you cannot see in top. flock(LOCK_EX) is the classic Linux advisory exclusive lock. Uncontended it is nearly free. Contended it becomes a serialized wait queue. Shared locks (LOCK_SH) are supposed to overlap. LOCK_NB is supposed to fail fast instead of sleeping.

This is a hands-on lab with measured numbers:

  1. Uncontended LOCK_EX acquire latency (microseconds).
  2. One holder, one waiter — block time ≈ hold time.
  3. Contended exclusive acquire p50 as worker count grows.
  4. LOCK_NB fail rate under contention.
  5. LOCK_SH wall-time overlap vs exclusive serialize.

Related links:

  • Pipe vs tmpfile IPC localhost lab
  • ulimit soft vs hard file descriptors lab
  • mmap vs read (+ O_DIRECT) localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 fcntl.flock on a local file (run/lab.lock). ProcessPool waiters. No NFS. No Docker. No GPU. No API keys. Affiliates: 0.

Verdict up front: uncontended p50 ≈ 0.33 µs; waiter behind a 50 ms hold ≈ 50.4 ms; contended w=8 p50 ≈ 36 ms (5 ms holds); LOCK_NB w=8 fail ≈ 99%; LOCK_SH wall ~6.7× vs serial hold.


What we compared

ArmAPIMetric
ALOCK_EX aloneacquire µs
Bholder + waiterwaiter block ms ≈ hold
CN processes LOCK_EXacquire p50 under queueing
DLOCK_EX|LOCK_NBfail %
ELOCK_SH vs LOCK_EXwall overlap vs serialize

Related links:

  • flock(2) — Linux man-pages

Lab topology

N worker processes → same lockfile (fcntl.flock)
Arm A: single process acquire/release × 300
Arm B: holder sleeps 50 ms under LOCK_EX; waiter blocks
Arm C: 2/4/8 workers × 25 rounds, hold=5 ms
Arm D: LOCK_NB attempts (w=1 vs w=8)
Arm E: 8× LOCK_SH vs 8× LOCK_EX (15 rounds × 10 ms hold)

Arm A — uncontended LOCK_EX

MetricValue
p500.33 µs
p950.35 µs
mean0.35 µs

Without a competitor, exclusive flock is noise-level cheap on a local filesystem.


Arm B — holder blocks waiter (hold = 50 ms)

MetricWaiter acquire
p5050.4 ms
p9552.6 ms
mean50.8 ms

The waiter slept for essentially the holder’s critical section. Contended flock is a queue, not a spin that burns CPU in our process (the waiter is blocked in the kernel).


Arm C — contended LOCK_EX (hold = 5 ms, 25 rounds/worker)

WorkersAcquire p50 msp95 msmean ms
24.977.134.90
415.220.615.3
835.944.735.1

More waiters → longer average queue. Rough intuition: with hold H and W workers, waits climb toward ~(W−1)·H under steady contention (8 workers × 5 ms → tens of ms — matches the table).


Arm D — LOCK_NB fail rate

WorkersGotFailFail %
16000%
8447699.2%

Non-blocking exclusive locks almost never succeed once eight processes hammer an 8 ms critical section. Use LOCK_NB when fail-fast / try-later is the product behavior — not when you need eventual entry.


Arm E — LOCK_SH overlap vs LOCK_EX serialize

8 processes × 15 rounds × 10 ms hold (ideal serial hold = 1.20 s):

ModeAcquire p50Wall svs ideal serial
LOCK_SH1.92 µs0.178~6.7× faster wall
LOCK_EX71.0 ms1.24~1.03× ideal

Shared locks overlapped almost to core count. Exclusive locks serialized to the hold budget. That is the whole design choice in one table.

Related links:

  • Pipe vs tmpfile IPC localhost lab

When exclusive still wins the design

  • Mutating a shared file / PID file / deploy stamp — readers must not see torn state.
  • Single-flight cron / job guards — one winner, others exit (LOCK_NB) or wait.
  • Not a distributed lock — this lab is local flock, not Redis/etcd/NFS edge cases.

Use LOCK_SH when many readers can safely overlap; measure wall time — Arm E is the template.


Pitfalls we hit (or avoided)

  1. Treating uncontended µs as the production number — Arm C is the one that hurts.
  2. Expecting LOCK_NB to “mostly work” under load — 99% fail at w=8.
  3. Assuming flock on NFS matches local — we did not test NFS; do not extrapolate.
  4. Forgetting advisory means cooperative — a process that never locks is unconstrained.
  5. Calling this a mutex benchmark for threads in one process — these are process waiters on a file.

Practical checklist

  • Critical section duration × waiter count ≈ user-visible stall — budget hold time first.
  • Fail-fast paths: LOCK_NB + retry/jitter; do not spin forever in userspace.
  • Reader-heavy: prefer LOCK_SH (or a different shared-state design); prove with wall overlap.
  • Report hold ms + worker count + p50 wait + NB fail% with every flock claim.
  • Keep lockfiles on local disk for this semantic; validate separately before trusting NFS.

Verdict

Local fcntl.flock: uncontended LOCK_EX p50 ≈ 0.33 µs; a 50 ms holder blocked the waiter ≈ 50.4 ms; under 5 ms holds, w=8 acquire p50 ≈ 36 ms; LOCK_NB failed ~99% at w=8; LOCK_SH delivered ~6.7× wall speedup vs serial while LOCK_EX matched ideal serialize (~1.03×).

Evidence path on the lab box: lab-evidence/27-flock-contention/results/. Affiliates: 0.

flock lock_exfile lock contentionfcntl flocklock_nbshared vs exclusive locklocalhost labsreprocess synchronization

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5 fcntl.flock on local file. Uncontended LOCK_EX p50 0.33 µs. Holder-block waiter hold=50 ms → waiter p50 50.4 ms. Contended LOCK_EX hold=5 ms: w2 p50 5.0 ms; w4 15.2 ms; w8 35.9 ms. LOCK_NB: w1 fail 0%; w8 fail 99.2%. LOCK_SH 8×15×10 ms: wall 0.178 s (~6.7× vs 1.2 s serial); LOCK_EX wall 1.24 s ≈ ideal. No Docker. Affiliates: 0. Evidence: lab-evidence/27-flock-contention/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — uncontended LOCK\_EX
  5. Arm B — holder blocks waiter (hold = 50 ms)
  6. Arm C — contended LOCK\_EX (hold = 5 ms, 25 rounds/worker)
  7. Arm D — LOCK\_NB fail rate
  8. Arm E — LOCK\_SH overlap vs LOCK\_EX serialize
  9. When exclusive still wins the design
  10. Pitfalls we hit (or avoided)
  11. Practical checklist
  12. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove