Plate 51
TCP_NODELAY vs Nagle: Localhost Lab with Real Numbers
Hands-on TCP_NODELAY vs Nagle: 1-byte reuse 124k vs 73k RPS; 32×16B writes favor Nagle (~1.7× fewer OutSegs). Delayed-ACK mean trap on localhost.
Aditya Challa6 min read
On this page
- Intro — what this post promises
- What Nagle and TCP\_NODELAY actually change
- Lab topology
- Arm 1 — reuse echo (the “RPC” shape)
- Arm 2 — many small writes, then one read (coalesce)
- Arm 3 — write-write-read + QUICKACK off
- How this sits next to keep-alive and UDS
- Pitfalls we hit (or avoided)
- Practical checklist
- Verdict
Intro — what this post promises
TCP_NODELAY disables Nagle’s algorithm. Folklore says “always turn it on for low latency.” That is half-true — and half-dangerous — once you measure what the sender is actually doing.
This is a localhost lab with measured numbers:
- What
getsockopt(TCP_NODELAY)reports on live sockets. - Tiny reuse echo: where NODELAY wins RPS.
- Many small writes then one read: where Nagle coalesces and wins.
- Write-write-read under
TCP_QUICKACK=0: mean ≫ p50 when Nagle meets delayed ACK. - A
TCP_CORKcontrol that looks like deliberate coalesce.
Related links:
- Unix Domain Socket vs TCP localhost lab
- HTTP Keep-Alive vs Connection: close lab
- Why your average latency graph is lying (p50 / p95 / p99)
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 TCP echo on 127.0.0.1:18310 (Nagle) and :18311 (TCP_NODELAY on accepted sockets). No Docker. No public bind. Affiliates: 0.
Verdict up front: reuse 1-byte echo favored NODELAY (~124k vs 73k~~ RPS). Batched 32×16 B writes favored Nagle (~~1.7× fewer OutSegs, ~2× higher RPS). With TCP_QUICKACK=0, Nagle write-write-read mean 119 µs vs p50 23 µs — the delayed-ACK tail — while NODELAY stayed at mean 22 µs.
What Nagle and TCP_NODELAY actually change
Nagle delays sending a small segment while an unacknowledged small segment is in flight, hoping to coalesce. TCP_NODELAY turns that off so each send can become a segment sooner.
Related links:
| Pattern | What we saw on loopback |
|---|---|
| One small request, one response (reuse) | NODELAY lower latency / higher RPS |
| Many tiny writes before a read | Nagle fewer segments, higher RPS |
| Two writes then read + delayed ACK pressure | Nagle mean blown by tail; NODELAY stable |
| Connect-per request | Handshake noise dominates; knob hard to see |
Lab topology
getsockopt check: NODELAY arm returned 1, Nagle arm returned 0.
Arm 1 — reuse echo (the “RPC” shape)
Single TCP connection, send payload, read echo. n=3000, warmup=150.
| Size | Arm | mean | p50 | p95 | RPS |
|---|---|---|---|---|---|
| 1 B | Nagle | 13.6 µs | 11.9 | 19.8 | ~73.4k |
| 1 B | NODELAY | 8.1 µs | 7.6 | 10.2 | ~124.1k |
| 64 B | Nagle | 10.6 µs | 11.3 | 12.8 | ~94.7k |
| 64 B | NODELAY | 9.8 µs | 9.2 | 10.8 | ~102.2k |
For chatty reuse protocols (game ticks, fine-grained RPC), NODELAY’s ~1.7× RPS on 1-byte messages is the intended win.
Arm 2 — many small writes, then one read (coalesce)
32 writes of 16 B (512 B total), then drain the echo.
| Arm | mean | RPS | OutSegs Δ (800 rounds) |
|---|---|---|---|
| Nagle | 41 µs | ~24.4k | 14202 |
| NODELAY | 79 µs | ~12.6k | 24296 (~1.71×) |
| TCP_CORK then uncork (control) | 38 µs | ~26.2k | (not counted) |
NODELAY was slower here because it emitted more segments. Nagle (and cork) batched. If your app already buffers, NODELAY is fine; if it write()s tiny pieces, Nagle or application-level buffering / cork saves work.
Arm 3 — write-write-read + QUICKACK off
Two writes totaling 64 B, then one read. Client set TCP_QUICKACK=0 to stress delayed ACK (best-effort on this stack).
| Arm | mean | p50 | p95 | RPS |
|---|---|---|---|---|
| Nagle | 118.8 µs | 23.4 | 44.6 | ~8.4k |
| NODELAY | 22.3 µs | 17.9 | 27.5 | ~44.9k |
This is why averages lie: Nagle’s p50 still looked fine while mean blew up. Pair with the percentiles playbook before you “fix” latency with one knob.
Related links:
Honesty: under default quickack on this loopback we did not see a classic ~40 ms stall. The trap still showed as a mean/p50 gap once quickack was forced off.
How this sits next to keep-alive and UDS
Keep-alive removes connect cost (our connect-per arms were ~20–27 µs of handshake noise). UDS removes TCP entirely for same-host IPC. TCP_NODELAY only changes TCP sender coalesce policy. Three different layers; three different graphs.
Related links:
Pitfalls we hit (or avoided)
- “Always set TCP_NODELAY” — wrong for write-heavy tiny-chunk apps without app buffering.
- Measuring only connect-per — handshake hides the knob.
- Looking only at mean — Nagle+delayed ACK is a tail/mean story.
- Assuming WAN 40 ms will appear on localhost — it did not under default quickack here.
- Forgetting both ends — server echo path also Nagles unless you set NODELAY on accepted sockets.
Practical checklist
- Know whether the app issues many small
sends or one buffered write. - For latency-sensitive reuse RPC: try
TCP_NODELAYand measure p50/p95, not only mean. - For throughput of tiny writes: prefer app buffering /
TCP_CORKover blind NODELAY. - Confirm with
getsockoptand, if needed, segment counters (/proc/net/snmpOutSegs). - Same-host and you care about µs: also consider UDS (separate lab).
Verdict
TCP_NODELAY is not a free latency cheat code. On this box it won chatty 1-byte reuse (~124k vs 73k~~ RPS) and lost batched tiny writes (~~1.7× more segments, ~half the RPS). Under delayed-ACK pressure, Nagle’s mean lied while p50 looked calm. Measure the write pattern you actually ship.
Evidence path on the lab box: lab-evidence/18-tcp-nodelay/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5 echo 127.0.0.1:18310 (Nagle) / :18311 (NODELAY). getsockopt on=1 off=0. Reuse 1B: Nagle mean 13.6µs ~73.4k RPS; NODELAY 8.1µs ~124.1k. Multi 32×16B segs arm: Nagle 41µs ~24.4k OutSegsΔ=14202; NODELAY 79µs ~12.6k OutSegsΔ=24296. WWR+QUICKACK=0: Nagle mean 118.8µs p50 23.4; NODELAY mean 22.3µs p50 17.9. Cork control ~38µs. No classic 40ms stall on default loopback quickack. Affiliates: 0. Evidence: lab-evidence/18-tcp-nodelay/.
Related links
Plate 42
TCP_QUICKACK Lab: Delayed ACK, Set-Once, and the Reassert Trap
Hands-on TCP_QUICKACK lab: default QA=1 on loopback; WWR set-once ~18 µs; reassert QA=1 under Nagle p95 ~48 ms, OutSegs ~13×. Real numbers, no Docker.
30 Sept 2026
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026