ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 57

  1. Blog
  2. /Observability & SRE

Context-Switch Microbench: Pipe Ping-Pong Process vs Thread

Hands-on ctx-switch lab: process pipe RTT p50 3.2 us (~0.50M sw/s); thread pipe 2.85 us; Event 9.3 us; 6 hogs blow p95 to 19 us. Real numbers, no Docker.

Aditya Challa·30 September 2026·5 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A/B — pipe ping-pong
  5. Arm C/D — Event ping-pong
  6. Arm E — yield and userspace floor
  7. Arm F — CPU interference (why p95 matters)
  8. How to read these numbers
  9. Pitfalls we hit (or avoided)
  10. Practical checklist
  11. Verdict

Intro — what this post promises

A context switch is what you pay when the kernel (or a userspace scheduler) stops one runnable task and starts another. Folklore quotes “a few microseconds.” This lab forces switches with a pipe ping-pong and measures round-trip latency.

This is a hands-on lab with measured numbers:

  1. Process↔process pipe ping-pong RTT and implied switch rate.
  2. Thread↔thread pipe ping-pong (same pattern, one address space).
  3. threading.Event vs asyncio.Event ping-pong.
  4. sched_yield and a userspace control.
  5. CPU hogs: p50 stays flat while p95/mean blow out.

Related links:

  • Pipe vs tmpfile IPC localhost lab
  • ProcessPool vs ThreadPool GIL localhost lab
  • epoll vs select FD_SETSIZE localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 pipes + events. No Docker. No GPU. No API keys. Affiliates: 0. This is a Python microbench, not lmbench C numbers.

Verdict up front: process pipe RTT p50 ≈ 3.2 µs (~0.50M sw/s if 2 switches/RTT); thread pipe ≈ 2.85 µs; threading.Event ≈ 9.3 µs; asyncio.Event ≈ 5.3 µs; under 6 CPU hogs, p50 stayed ~3.6 µs while p95 ≈ 19 µs.


What we compared

ArmMechanismWhat it forces
A2 processes, 2 pipeskernel block/wake across address spaces
B2 threads, 2 pipeskernel block/wake, shared process
Cthreading.Eventfutex-style wait/wake
Dasyncio.Eventcooperative task switch (usually no kernel ctx)
Esched_yield / userspace bytelower bounds / noise floor
FArm A + CPU hogsscheduling interference on tails

Related links:

  • pipe(7) — Linux man-pages
  • sched_yield(2)

Lab topology

Parent writes 1 byte on pipe1 → child blocks-in-read then writes pipe2 → parent reads
Timed unit = one RTT (typically ~2 forced switches)
Also: Event ping-pong; sched_yield loop; 0/2/6 busy-loop hog processes

Assumption we label explicitly: 1 RTT ≈ 2 switches for the pipe arms. Rates are approximate.


Arm A/B — pipe ping-pong

n=20 000 timed RTTs (after warmup):

ArmRTT p50p95meanRTT RPS~M sw/s
process pipe3.21 µs5.444.02~249k~0.50
thread pipe2.85 µs4.563.52~284k~0.57

Same pattern; threads were ~13% faster p50 on this box (no TLB/mm switch). Stability×3 median p50: process 3.35 µs, thread 2.84 µs.

Child /proc/<pid>/status voluntary switches ≈ 0.8/op on the interference runs — the wakeups are real, not imaginary.

Related links:

  • Pipe vs tmpfile IPC localhost lab

Arm C/D — Event ping-pong

ArmRTT p50p95meannote
threading.Event9.28 µs17.811.0kernel wait/wake
asyncio.Event5.33 µs6.375.83cooperative; typically no kernel ctx switch

threading.Event was the slowest ping-pong here — heavier than the pipe path. Asyncio was mid-pack: cooperative ≠ free; you still pay Python scheduling even when the kernel stays out.


Arm E — yield and userspace floor

Opp50
os.sched_yield() (single thread)0.34 µs
userspace byte assign control0.09 µs

A lone sched_yield is not a two-task switch benchmark — often nothing else is runnable. Treat it as a syscall floor, not “context switch cost.”


Arm F — CPU interference (why p95 matters)

Process pipe RTT with busy-loop hogs:

Hogsp50 µsp95 µsmean µs~M sw/s
03.516.674.160.48
24.135.064.830.41
63.5918.86.260.32

Median barely moved. p95 and mean took the hit — the classic “average latency graph is lying” failure mode under load.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

How to read these numbers

  • Pipe RTT is a forced switch microbench, not your app’s p99.
  • Process vs thread here is close; the bigger jump was Event vs pipe.
  • Asyncio Event avoids kernel ctx switches but still costs microseconds in CPython.
  • Under CPU contention, publish p95/p99, not only p50.

Pitfalls we hit (or avoided)

  1. Calling sched_yield a context-switch benchmark — needs a competing runnable task.
  2. Using only parent /proc/self voluntary counts — child’s switches live under the child pid.
  3. Reporting M sw/s without the 2-switches/RTT assumption — we label it.
  4. Comparing to C lmbench headlines — different runtime; do not paste those as ours.
  5. Trusting p50 under load — Arm F’s p95 is the story.

Practical checklist

  • When claiming switch cost, show method (pipe/Event) + p50/p95 + core count.
  • Separate kernel switches from userspace scheduler hops (asyncio).
  • Under shared hosts, watch tails when runnable count rises.
  • Prefer fewer blocking hops on latency-critical paths; batch work per wake.
  • Keep microbench evidence next to any “µs context switch” slide.

Verdict

Python pipe ping-pong on this box: process RTT p50 ≈ 3.2 µs (~0.50M sw/s); thread pipe ≈ 2.85 µs; threading.Event ≈ 9.3 µs; asyncio.Event ≈ 5.3 µs. With 6 CPU hogs, p50 stayed ~3.6 µs while p95 climbed to ~19 µs.

Evidence path on the lab box: lab-evidence/29-context-switch/results/. Affiliates: 0.

context switch benchmarkpipe ping-pong latencyprocess vs thread switchsched_yield costvoluntary ctxt switcheslocalhost labsrelinux scheduling

Lab evidence

What I found running this

Lab 1 Oct 2026 IST on a shared Linux box (8 cores, kernel 6.12), Python 3.13.5. I ran 20,000 process and thread pipe ping-pong RTTs, threading.Event and asyncio.Event, sched_yield, userspace control, and CPU-hog interference. Process pipe p50 3.21 us; thread pipe 2.85 us; threading.Event 9.28 us; asyncio.Event 5.33 us. With 6 CPU hogs, p50 stayed 3.59 us but p95 rose to 18.8 us. No Docker, GPU, or API keys; affiliates 0.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A/B — pipe ping-pong
  5. Arm C/D — Event ping-pong
  6. Arm E — yield and userspace floor
  7. Arm F — CPU interference (why p95 matters)
  8. How to read these numbers
  9. Pitfalls we hit (or avoided)
  10. Practical checklist
  11. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove