Plate 60
Pipe vs Tmpfile IPC: Localhost Throughput and RTT Lab
Hands-on pipe vs tmpfile IPC lab: 5.32 vs 1.51 GB/s bulk (fsync 0.57); 64 B RTT p50 3.3 us vs 70 us vs 289 us. Localhost measured, no Docker.
Aditya Challa6 min read
Intro — what this post promises
Need to move bytes between processes on one box? Two defaults fight for the job: pipes (os.pipe / shell |) and temporary files. Folklore says “files are fine on SSD.” This lab measures both shapes with the same payload sizes.
This is a hands-on lab with measured numbers:
- Bulk throughput: 64 MiB through a pipe vs write+read of a tmpfile (with and without
fsync). - Request/response RTT: duplex pipe (fork parent↔child) vs a file “mailbox” open/write/read loop.
- How chunk size and
fsyncchange the story. - When a tmpfile is still the right tool despite losing the race.
Related links:
- Unix Domain Socket vs TCP localhost lab
- HTTP Keep-Alive vs Connection: close lab
- Why your average latency graph is lying (p50 / p95 / p99)
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. Pipes via os.pipe + fork; tmpfiles under /workspace (overlay). No Docker. No GPU. Affiliates: 0.
Verdict up front: best pipe bulk was 5.32 GB/s (64 KiB chunks). Same 64 MiB via tmpfile buffered hit 1.51 GB/s (3.5×~~ slower); with fsync 0.57 GB/s (~~9×). RTT for 64 B: pipe p50 3.3 µs (~285k RPS) vs tmpfile 70 µs (~12k) vs tmpfile+fsync 289 µs (~2.5k).
What we compared
| Medium | Shape in this lab |
|---|---|
| Pipe | Producer forks; parent writes, child reads (throughput) or duplex request/response (RTT) |
| Tmpfile buffered | Write whole object, then read it back (page cache friendly) |
Tmpfile + fsync | Same, but flush + os.fsync after write (durability tax) |
Related links:
Honesty on RTT: the tmpfile arm is a same-process mailbox (open/write/read). That is an optimistic lower bound for “leave a file for the next stage.” The pipe RTT is a true two-process duplex. If anything, real multi-process file polling would be worse for files.
Lab topology
Arm A — bulk throughput (64 MiB)
| Medium | Chunk | GB/s | Seconds |
|---|---|---|---|
| Pipe | 4 KiB | 1.88 | 0.036 |
| Pipe | 64 KiB | 5.32 | 0.013 |
| Pipe | 1 MiB | 1.41 | 0.048 |
| Tmpfile buffered | 4 KiB | 0.93 | 0.072 |
| Tmpfile buffered | 64 KiB | 1.51 | 0.044 |
| Tmpfile buffered | 1 MiB | 1.26 | 0.053 |
| Tmpfile + fsync | 4 KiB | 0.44 | 0.153 |
| Tmpfile + fsync | 64 KiB | 0.57 | 0.118 |
| Tmpfile + fsync | 1 MiB | 0.63 | 0.107 |
Pipe peaked at 64 KiB chunks. Tmpfile never caught pipe on this pass; fsync roughly halved to thirded buffered tmpfile GB/s. A 128 MiB confirm pass kept the same ranking (pipe ~3.7 GB/s at 64 KiB vs tmpfile ~1.4 / fsync ~0.54).
Arm B — request/response RTT
| Size | Medium | p50 | mean | RPS |
|---|---|---|---|---|
| 64 B | Pipe duplex | 3.3 µs | 3.5 | ~285k |
| 64 B | Tmpfile | 70.5 µs | 83.6 | ~12.0k |
| 64 B | Tmpfile + fsync | 288.7 µs | 396 | ~2.5k |
| 1 B | Pipe | 5.2 µs | 5.8 | ~172k |
| 1 B | Tmpfile | 70.0 µs | 76.2 | ~13.1k |
| 1 B | Tmpfile + fsync | 285 µs | 384 | ~2.6k |
| 64 KiB | Pipe | 18.2 µs | 20.1 | ~50k |
| 64 KiB | Tmpfile | 185 µs | 202 | ~5.0k |
| 64 KiB | Tmpfile + fsync | 471 µs | 633 | ~1.6k |
For chatty handoffs, pipe’s ~21× p50 advantage over buffered files (and ~87× over fsync) is the operational punchline. Durability is not free.
When tmpfiles still win the design
Pipes lose data if nobody is reading and the buffer fills; they do not survive process restart; they are awkward for random access and for “many consumers later.” Tmpfiles (or real paths) win when you need replay, crash durability, multi-reader, or a simple ops story (ls the artifact). This lab measures speed, not those product constraints.
UDS sits in between for same-host RPC — see the UDS vs TCP lab — while pipes stay the cheapest streaming primitive inside one machine.
Related links:
Pitfalls we hit (or avoided)
- Comparing pipe RTT to an unfair file poller — we used a same-process mailbox and said so; real watchers cost more.
- Forgetting
fsync— buffered GB/s looks fine until durability is required. - Wrong chunk size — pipe at 1 MiB was slower than 64 KiB here; measure your chunk.
- Calling this shared-memory or
mmap— different tools; not in this package. - Assuming NVMe makes files “as fast as pipes” — not for RTT-shaped IPC on this box.
Practical checklist
- Streaming one producer → one consumer on one host: start with a pipe (or UDS if you need sockets).
- Need durability or multi-stage replay: tmpfile/object store — budget the
fsynctax. - Measure p50/p95 for chatty IPC; mean alone hides fsync spikes.
- Tune chunk size; do not assume bigger is always better.
- Do not confuse this with network RPC or cross-host queues.
Verdict
Pipes crushed tmpfiles for both bulk and RTT on this box: 5.32 GB/s vs 1.51 GB/s buffered (64 MiB / 64 KiB chunks), and 3.3 µs vs 70 µs p50 for 64 B RTT — 289 µs once fsync entered. Use files when you need persistence or ops visibility; use pipes when you need streaming speed between live processes.
Evidence path on the lab box: lab-evidence/19-pipe-vs-tmpfile/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5. 64 MiB throughput: pipe chunk 64 KiB 5.322 GB/s; tmpfile buffered 1.511 GB/s; tmpfile+fsync 0.569 GB/s. RTT 64 B: pipe duplex p50 3.3 µs ~285k RPS; tmpfile mailbox p50 70.5 µs ~12k; +fsync p50 288.7 µs ~2.5k. Pipe 1 B p50 5.2 µs; 64 KiB p50 18.2 µs. 128 MiB confirm: pipe ~3.74 GB/s vs tmpfile ~1.44 / fsync ~0.54. Affiliates: 0. Evidence: lab-evidence/19-pipe-vs-tmpfile/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026