ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 60

  1. Blog
  2. /Observability & SRE

Pipe vs Tmpfile IPC: Localhost Throughput and RTT Lab

Hands-on pipe vs tmpfile IPC lab: 5.32 vs 1.51 GB/s bulk (fsync 0.57); 64 B RTT p50 3.3 us vs 70 us vs 289 us. Localhost measured, no Docker.

Aditya Challa·30 September 2026·6 min read

Lab
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — bulk throughput (64 MiB)
  5. Arm B — request/response RTT
  6. When tmpfiles still win the design
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict

Intro — what this post promises

Need to move bytes between processes on one box? Two defaults fight for the job: pipes (os.pipe / shell |) and temporary files. Folklore says “files are fine on SSD.” This lab measures both shapes with the same payload sizes.

This is a hands-on lab with measured numbers:

  1. Bulk throughput: 64 MiB through a pipe vs write+read of a tmpfile (with and without fsync).
  2. Request/response RTT: duplex pipe (fork parent↔child) vs a file “mailbox” open/write/read loop.
  3. How chunk size and fsync change the story.
  4. When a tmpfile is still the right tool despite losing the race.

Related links:

  • Unix Domain Socket vs TCP localhost lab
  • HTTP Keep-Alive vs Connection: close lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. Pipes via os.pipe + fork; tmpfiles under /workspace (overlay). No Docker. No GPU. Affiliates: 0.

Verdict up front: best pipe bulk was 5.32 GB/s (64 KiB chunks). Same 64 MiB via tmpfile buffered hit 1.51 GB/s (3.5×~~ slower); with fsync 0.57 GB/s (~~9×). RTT for 64 B: pipe p50 3.3 µs (~285k RPS) vs tmpfile 70 µs (~12k) vs tmpfile+fsync 289 µs (~2.5k).


What we compared

MediumShape in this lab
PipeProducer forks; parent writes, child reads (throughput) or duplex request/response (RTT)
Tmpfile bufferedWrite whole object, then read it back (page cache friendly)
Tmpfile + fsyncSame, but flush + os.fsync after write (durability tax)

Related links:

  • man 2 pipe
  • man 2 fsync

Honesty on RTT: the tmpfile arm is a same-process mailbox (open/write/read). That is an optimistic lower bound for “leave a file for the next stage.” The pipe RTT is a true two-process duplex. If anything, real multi-process file polling would be worse for files.


Lab topology

Throughput: 64 MiB  →  pipe(chunk)  |  tmpfile write→read  |  tmpfile+fsync
RTT:        sizes 1 / 64 / 1K / 64K  →  pipe duplex  |  file mailbox ± fsync
Confirm:    128 MiB throughput pass (same order of magnitude)

Arm A — bulk throughput (64 MiB)

MediumChunkGB/sSeconds
Pipe4 KiB1.880.036
Pipe64 KiB5.320.013
Pipe1 MiB1.410.048
Tmpfile buffered4 KiB0.930.072
Tmpfile buffered64 KiB1.510.044
Tmpfile buffered1 MiB1.260.053
Tmpfile + fsync4 KiB0.440.153
Tmpfile + fsync64 KiB0.570.118
Tmpfile + fsync1 MiB0.630.107

Pipe peaked at 64 KiB chunks. Tmpfile never caught pipe on this pass; fsync roughly halved to thirded buffered tmpfile GB/s. A 128 MiB confirm pass kept the same ranking (pipe ~3.7 GB/s at 64 KiB vs tmpfile ~1.4 / fsync ~0.54).


Arm B — request/response RTT

SizeMediump50meanRPS
64 BPipe duplex3.3 µs3.5~285k
64 BTmpfile70.5 µs83.6~12.0k
64 BTmpfile + fsync288.7 µs396~2.5k
1 BPipe5.2 µs5.8~172k
1 BTmpfile70.0 µs76.2~13.1k
1 BTmpfile + fsync285 µs384~2.6k
64 KiBPipe18.2 µs20.1~50k
64 KiBTmpfile185 µs202~5.0k
64 KiBTmpfile + fsync471 µs633~1.6k

For chatty handoffs, pipe’s ~21× p50 advantage over buffered files (and ~87× over fsync) is the operational punchline. Durability is not free.


When tmpfiles still win the design

Pipes lose data if nobody is reading and the buffer fills; they do not survive process restart; they are awkward for random access and for “many consumers later.” Tmpfiles (or real paths) win when you need replay, crash durability, multi-reader, or a simple ops story (ls the artifact). This lab measures speed, not those product constraints.

UDS sits in between for same-host RPC — see the UDS vs TCP lab — while pipes stay the cheapest streaming primitive inside one machine.

Related links:

  • Unix Domain Socket vs TCP localhost lab

Pitfalls we hit (or avoided)

  1. Comparing pipe RTT to an unfair file poller — we used a same-process mailbox and said so; real watchers cost more.
  2. Forgetting fsync — buffered GB/s looks fine until durability is required.
  3. Wrong chunk size — pipe at 1 MiB was slower than 64 KiB here; measure your chunk.
  4. Calling this shared-memory or mmap — different tools; not in this package.
  5. Assuming NVMe makes files “as fast as pipes” — not for RTT-shaped IPC on this box.

Practical checklist

  • Streaming one producer → one consumer on one host: start with a pipe (or UDS if you need sockets).
  • Need durability or multi-stage replay: tmpfile/object store — budget the fsync tax.
  • Measure p50/p95 for chatty IPC; mean alone hides fsync spikes.
  • Tune chunk size; do not assume bigger is always better.
  • Do not confuse this with network RPC or cross-host queues.

Verdict

Pipes crushed tmpfiles for both bulk and RTT on this box: 5.32 GB/s vs 1.51 GB/s buffered (64 MiB / 64 KiB chunks), and 3.3 µs vs 70 µs p50 for 64 B RTT — 289 µs once fsync entered. Use files when you need persistence or ops visibility; use pipes when you need streaming speed between live processes.

Evidence path on the lab box: lab-evidence/19-pipe-vs-tmpfile/results/. Affiliates: 0.

pipe vs file ipcos.pipetmpfile performancefsync costinterprocess communicationlocalhost labsrelinux pipes

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5. 64 MiB throughput: pipe chunk 64 KiB 5.322 GB/s; tmpfile buffered 1.511 GB/s; tmpfile+fsync 0.569 GB/s. RTT 64 B: pipe duplex p50 3.3 µs ~285k RPS; tmpfile mailbox p50 70.5 µs ~12k; +fsync p50 288.7 µs ~2.5k. Pipe 1 B p50 5.2 µs; 64 KiB p50 18.2 µs. 128 MiB confirm: pipe ~3.74 GB/s vs tmpfile ~1.44 / fsync ~0.54. Affiliates: 0. Evidence: lab-evidence/19-pipe-vs-tmpfile/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — bulk throughput (64 MiB)
  5. Arm B — request/response RTT
  6. When tmpfiles still win the design
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove