Plate 64
Unix Domain Socket vs TCP Localhost: Latency and Throughput Lab
Hands-on UDS vs TCP loopback echo: connect-per ~3× RPS (49k vs 16k); reuse ~1.3×. Real Python numbers, no Docker.
Aditya Challa6 min read
Intro — what this post promises
Same-host services often talk over 127.0.0.1:PORT out of habit. A Unix domain socket (AF_UNIX) skips the TCP/IP stack for that hop. This post measures both on one box.
This is a localhost lab with measured numbers:
- A fair echo comparison: TCP loopback vs Unix domain socket.
- Two client modes: new connect every request vs one socket, many round-trips.
- Three payload sizes: 64 B, 4 KiB, 64 KiB.
- When the gap is huge (
3×) and when it shrinks (25–30%). - What this does not claim (WAN, TLS, HTTP/2).
Related links:
- HTTP Keep-Alive vs Connection: close lab
- Nginx limit_req rate-limit lab
- Why your average latency graph is lying (p50 / p95 / p99)
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). Python 3.13.5 echo servers — TCP on 127.0.0.1:18210 only, UDS at a workspace path with mode 0600. Client used TCP_NODELAY on TCP. Warmup 100 ops discarded per arm. No Docker. No public bind. Affiliates: 0.
Verdict up front: for same-host IPC, UDS wins. The win is largest when you connect a lot; persistent connections shrink it. Prefer UDS (or a well-tuned keep-alive TCP client) when both ends live on one machine.
What UDS is (and is not)
A Unix domain socket is a filesystem-path (or abstract-name) endpoint for stream or datagram IPC on one host. No IP addresses, no TCP handshake, no TIME_WAIT from TCP.
Related links:
Mental model that matched our lab:
| Choice | What it did here |
|---|---|
TCP 127.0.0.1 | Full TCP stack on loopback; still cheap, still a handshake per connect |
| UDS path socket | Kernel IPC; connect is cheaper; no TCP state machine |
| Connect-per-op | Pays connect + teardown every round-trip — amplifies UDS advantage |
| Single-socket reuse | Amortizes connect; gap shrinks to throughput/copy differences |
TCP_NODELAY | Avoids Nagle delay on tiny TCP messages (fair small-payload arm) |
UDS is not a substitute for auth across machines, not HTTP, and not a fix for a slow accept loop (see the backlog lab).
Lab topology
Minimal shapes:
Client arms: send N fixed-size payloads, read the echo back, record wall time and per-op latency percentiles.
64-byte payload: connect tax dominates
n=5000 sequential echo round-trips, 64-byte payload:
| Transport | Mode | Wall | Approx RPS | Latency p50 |
|---|---|---|---|---|
| TCP | New connect each | 0.302 s | ~16531 | 0.021 ms |
| TCP | One socket, reuse | 0.042 s | ~118215 | 0.0075 ms |
| UDS | New connect each | 0.101 s | ~49374 | 0.0105 ms |
| UDS | One socket, reuse | 0.032 s | ~154621 | 0.0055 ms |
Connect-per: UDS delivered about 3.0× the RPS of TCP (~49374 / ~16531).
Reuse: UDS was about 1.31× TCP (~154621 / ~118215).
That is the headline: if your “localhost microservice” opens a fresh TCP connection per RPC, you are leaving a large same-host win on the table — either switch to UDS or reuse connections (same lesson as the keep-alive lab, different layer).
4 KiB and 64 KiB: gap shrinks, UDS still ahead
4 KiB, n=3000:
| Transport | Mode | Approx RPS | Latency p50 |
|---|---|---|---|
| TCP | Connect-per | ~20484 | 0.019 ms |
| TCP | Reuse | ~115597 | 0.0080 ms |
| UDS | Connect-per | 0.012 ms | |
| UDS | Reuse | 0.0061 ms |
64 KiB, n=1000, reuse only (connect cost is noise next to copy):
| Transport | Approx RPS | Latency p50 |
|---|---|---|
| TCP reuse | ~34240 | 0.024 ms |
| UDS reuse | 0.020 ms |
Larger payloads move the bottleneck toward copying and scheduling. UDS still led; the dramatic 3× story is a connect-heavy story.
How to read this next to keep-alive
Keep-alive is about reusing HTTP/TCP across requests. UDS is about skipping TCP for same-host peers. They stack:
- Same host → prefer UDS (or a local abstract socket) when the stack allows it (postgres
host, redisunixsocket, nginxlisten unix:…, gRPC UDS, etc.). - Must use TCP → reuse connections; do not shell out to a new
curlper call. - Watch percentiles, not only mean RPS — our p50/p95/p99 columns are in the evidence JSON.
Related links:
Pitfalls we hit (or avoided)
- Comparing a reused UDS socket to a connect-per TCP client — unfair. We published both modes for both transports.
- Forgetting
TCP_NODELAYon tiny TCP messages — can inflate TCP latency via Nagle; we enabled it for the TCP client. - World-writable socket paths — we used mode
0600on the socket file. Treat the path like a secret capability. - Calling this a WAN result — it is not. Cross-host you need TCP/TLS; measure that separately.
- Ignoring accept backlog — a fast UDS client against a slow accept loop still drops. Pair with the listen-backlog lab.
Practical checklist
- Same-host DB/cache/proxy: check for a native Unix socket option before defaulting to
127.0.0.1. - If TCP localhost is mandatory: connection pool / keep-alive; measure connect-per vs reuse.
- Lock down socket filesystem permissions (
0600/ dedicated directory). - Report p50/p95, not only average RPS.
- Do not claim WAN or TLS wins from a loopback echo lab.
Verdict
On this box, Unix domain sockets beat TCP loopback for echo IPC. The ~3× RPS win on 64-byte connect-per arms is the clearest signal; reuse still favored UDS by roughly 25–30%. Use UDS for same-host hops when you can; when you cannot, stop paying a TCP handshake per call.
Evidence path on the lab box: lab-evidence/14-uds-vs-tcp/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5 echo servers: TCP 127.0.0.1:18210 (TCP_NODELAY client) and AF_UNIX path. Connect-per vs single-socket reuse. 64 B n=5000: TCP connect ~16531 rps p50 0.021 ms; TCP reuse ~118215 rps p50 0.0075 ms; UDS connect ~49374 rps p50 0.0105 ms (~3.0×); UDS reuse ~154621 rps p50 0.0055 ms (~1.31×). 4 KiB n=3000: UDS connect ~2.46× TCP connect; reuse ~1.29×. 64 KiB reuse n=1000: TCP ~34240 rps vs UDS ~42819 rps (~1.25×). Localhost only; not a WAN/TLS claim. Affiliates: 0. Evidence: lab-evidence/14-uds-vs-tcp/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026