Plate 57
Context-Switch Microbench: Pipe Ping-Pong Process vs Thread
Hands-on ctx-switch lab: process pipe RTT p50 3.2 us (~0.50M sw/s); thread pipe 2.85 us; Event 9.3 us; 6 hogs blow p95 to 19 us. Real numbers, no Docker.
Aditya Challa5 min read
Intro — what this post promises
A context switch is what you pay when the kernel (or a userspace scheduler) stops one runnable task and starts another. Folklore quotes “a few microseconds.” This lab forces switches with a pipe ping-pong and measures round-trip latency.
This is a hands-on lab with measured numbers:
- Process↔process pipe ping-pong RTT and implied switch rate.
- Thread↔thread pipe ping-pong (same pattern, one address space).
threading.Eventvsasyncio.Eventping-pong.sched_yieldand a userspace control.- CPU hogs: p50 stays flat while p95/mean blow out.
Related links:
- Pipe vs tmpfile IPC localhost lab
- ProcessPool vs ThreadPool GIL localhost lab
- epoll vs select FD_SETSIZE localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 pipes + events. No Docker. No GPU. No API keys. Affiliates: 0. This is a Python microbench, not lmbench C numbers.
Verdict up front: process pipe RTT p50 ≈ 3.2 µs (~0.50M sw/s if 2 switches/RTT); thread pipe ≈ 2.85 µs; threading.Event ≈ 9.3 µs; asyncio.Event ≈ 5.3 µs; under 6 CPU hogs, p50 stayed ~3.6 µs while p95 ≈ 19 µs.
What we compared
| Arm | Mechanism | What it forces |
|---|---|---|
| A | 2 processes, 2 pipes | kernel block/wake across address spaces |
| B | 2 threads, 2 pipes | kernel block/wake, shared process |
| C | threading.Event | futex-style wait/wake |
| D | asyncio.Event | cooperative task switch (usually no kernel ctx) |
| E | sched_yield / userspace byte | lower bounds / noise floor |
| F | Arm A + CPU hogs | scheduling interference on tails |
Related links:
Lab topology
Assumption we label explicitly: 1 RTT ≈ 2 switches for the pipe arms. Rates are approximate.
Arm A/B — pipe ping-pong
n=20 000 timed RTTs (after warmup):
| Arm | RTT p50 | p95 | mean | RTT RPS | ~M sw/s |
|---|---|---|---|---|---|
| process pipe | 3.21 µs | 5.44 | 4.02 | ~249k | ~0.50 |
| thread pipe | 2.85 µs | 4.56 | 3.52 | ~284k | ~0.57 |
Same pattern; threads were ~13% faster p50 on this box (no TLB/mm switch). Stability×3 median p50: process 3.35 µs, thread 2.84 µs.
Child /proc/<pid>/status voluntary switches ≈ 0.8/op on the interference runs — the wakeups are real, not imaginary.
Related links:
Arm C/D — Event ping-pong
| Arm | RTT p50 | p95 | mean | note |
|---|---|---|---|---|
threading.Event | 9.28 µs | 17.8 | 11.0 | kernel wait/wake |
asyncio.Event | 5.33 µs | 6.37 | 5.83 | cooperative; typically no kernel ctx switch |
threading.Event was the slowest ping-pong here — heavier than the pipe path. Asyncio was mid-pack: cooperative ≠ free; you still pay Python scheduling even when the kernel stays out.
Arm E — yield and userspace floor
| Op | p50 |
|---|---|
os.sched_yield() (single thread) | 0.34 µs |
| userspace byte assign control | 0.09 µs |
A lone sched_yield is not a two-task switch benchmark — often nothing else is runnable. Treat it as a syscall floor, not “context switch cost.”
Arm F — CPU interference (why p95 matters)
Process pipe RTT with busy-loop hogs:
| Hogs | p50 µs | p95 µs | mean µs | ~M sw/s |
|---|---|---|---|---|
| 0 | 3.51 | 6.67 | 4.16 | 0.48 |
| 2 | 4.13 | 5.06 | 4.83 | 0.41 |
| 6 | 3.59 | 18.8 | 6.26 | 0.32 |
Median barely moved. p95 and mean took the hit — the classic “average latency graph is lying” failure mode under load.
Related links:
How to read these numbers
- Pipe RTT is a forced switch microbench, not your app’s p99.
- Process vs thread here is close; the bigger jump was Event vs pipe.
- Asyncio Event avoids kernel ctx switches but still costs microseconds in CPython.
- Under CPU contention, publish p95/p99, not only p50.
Pitfalls we hit (or avoided)
- Calling
sched_yielda context-switch benchmark — needs a competing runnable task. - Using only parent
/proc/selfvoluntary counts — child’s switches live under the child pid. - Reporting M sw/s without the 2-switches/RTT assumption — we label it.
- Comparing to C lmbench headlines — different runtime; do not paste those as ours.
- Trusting p50 under load — Arm F’s p95 is the story.
Practical checklist
- When claiming switch cost, show method (pipe/Event) + p50/p95 + core count.
- Separate kernel switches from userspace scheduler hops (asyncio).
- Under shared hosts, watch tails when runnable count rises.
- Prefer fewer blocking hops on latency-critical paths; batch work per wake.
- Keep microbench evidence next to any “µs context switch” slide.
Verdict
Python pipe ping-pong on this box: process RTT p50 ≈ 3.2 µs (~0.50M sw/s); thread pipe ≈ 2.85 µs; threading.Event ≈ 9.3 µs; asyncio.Event ≈ 5.3 µs. With 6 CPU hogs, p50 stayed ~3.6 µs while p95 climbed to ~19 µs.
Evidence path on the lab box: lab-evidence/29-context-switch/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST on a shared Linux box (8 cores, kernel 6.12), Python 3.13.5. I ran 20,000 process and thread pipe ping-pong RTTs, threading.Event and asyncio.Event, sched_yield, userspace control, and CPU-hog interference. Process pipe p50 3.21 us; thread pipe 2.85 us; threading.Event 9.28 us; asyncio.Event 5.33 us. With 6 CPU hogs, p50 stayed 3.59 us but p95 rose to 18.8 us. No Docker, GPU, or API keys; affiliates 0.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026