ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 51

  1. Blog
  2. /Observability & SRE

TCP_NODELAY vs Nagle: Localhost Lab with Real Numbers

Hands-on TCP_NODELAY vs Nagle: 1-byte reuse 124k vs 73k RPS; 32×16B writes favor Nagle (~1.7× fewer OutSegs). Delayed-ACK mean trap on localhost.

Aditya Challa·30 September 2026·6 min read

Lab
On this page
  1. Intro — what this post promises
  2. What Nagle and TCP\_NODELAY actually change
  3. Lab topology
  4. Arm 1 — reuse echo (the “RPC” shape)
  5. Arm 2 — many small writes, then one read (coalesce)
  6. Arm 3 — write-write-read + QUICKACK off
  7. How this sits next to keep-alive and UDS
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict

Intro — what this post promises

TCP_NODELAY disables Nagle’s algorithm. Folklore says “always turn it on for low latency.” That is half-true — and half-dangerous — once you measure what the sender is actually doing.

This is a localhost lab with measured numbers:

  1. What getsockopt(TCP_NODELAY) reports on live sockets.
  2. Tiny reuse echo: where NODELAY wins RPS.
  3. Many small writes then one read: where Nagle coalesces and wins.
  4. Write-write-read under TCP_QUICKACK=0: mean ≫ p50 when Nagle meets delayed ACK.
  5. A TCP_CORK control that looks like deliberate coalesce.

Related links:

  • Unix Domain Socket vs TCP localhost lab
  • HTTP Keep-Alive vs Connection: close lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 TCP echo on 127.0.0.1:18310 (Nagle) and :18311 (TCP_NODELAY on accepted sockets). No Docker. No public bind. Affiliates: 0.

Verdict up front: reuse 1-byte echo favored NODELAY (~124k vs 73k~~ RPS). Batched 32×16 B writes favored Nagle (~~1.7× fewer OutSegs, ~2× higher RPS). With TCP_QUICKACK=0, Nagle write-write-read mean 119 µs vs p50 23 µs — the delayed-ACK tail — while NODELAY stayed at mean 22 µs.


What Nagle and TCP_NODELAY actually change

Nagle delays sending a small segment while an unacknowledged small segment is in flight, hoping to coalesce. TCP_NODELAY turns that off so each send can become a segment sooner.

Related links:

  • man 7 tcp — TCP_NODELAY
  • RFC 896 — Congestion Control in IP/TCP (Nagle)
PatternWhat we saw on loopback
One small request, one response (reuse)NODELAY lower latency / higher RPS
Many tiny writes before a readNagle fewer segments, higher RPS
Two writes then read + delayed ACK pressureNagle mean blown by tail; NODELAY stable
Connect-per requestHandshake noise dominates; knob hard to see

Lab topology

Server A :18310  accept sockets with Nagle (default)
Server B :18311  accept sockets with TCP_NODELAY=1
Client matches the arm (setsockopt TCP_NODELAY on/off)
Arms: reuse echo · connect-per · multi-write · WWR · cork · /proc/net/snmp OutSegsΔ

getsockopt check: NODELAY arm returned 1, Nagle arm returned 0.


Arm 1 — reuse echo (the “RPC” shape)

Single TCP connection, send payload, read echo. n=3000, warmup=150.

SizeArmmeanp50p95RPS
1 BNagle13.6 µs11.919.8~73.4k
1 BNODELAY8.1 µs7.610.2~124.1k
64 BNagle10.6 µs11.312.8~94.7k
64 BNODELAY9.8 µs9.210.8~102.2k

For chatty reuse protocols (game ticks, fine-grained RPC), NODELAY’s ~1.7× RPS on 1-byte messages is the intended win.


Arm 2 — many small writes, then one read (coalesce)

32 writes of 16 B (512 B total), then drain the echo.

ArmmeanRPSOutSegs Δ (800 rounds)
Nagle41 µs~24.4k14202
NODELAY79 µs~12.6k24296 (~1.71×)
TCP_CORK then uncork (control)38 µs~26.2k(not counted)

NODELAY was slower here because it emitted more segments. Nagle (and cork) batched. If your app already buffers, NODELAY is fine; if it write()s tiny pieces, Nagle or application-level buffering / cork saves work.


Arm 3 — write-write-read + QUICKACK off

Two writes totaling 64 B, then one read. Client set TCP_QUICKACK=0 to stress delayed ACK (best-effort on this stack).

Armmeanp50p95RPS
Nagle118.8 µs23.444.6~8.4k
NODELAY22.3 µs17.927.5~44.9k

This is why averages lie: Nagle’s p50 still looked fine while mean blew up. Pair with the percentiles playbook before you “fix” latency with one knob.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Honesty: under default quickack on this loopback we did not see a classic ~40 ms stall. The trap still showed as a mean/p50 gap once quickack was forced off.


How this sits next to keep-alive and UDS

Keep-alive removes connect cost (our connect-per arms were ~20–27 µs of handshake noise). UDS removes TCP entirely for same-host IPC. TCP_NODELAY only changes TCP sender coalesce policy. Three different layers; three different graphs.

Related links:

  • HTTP Keep-Alive vs Connection: close lab
  • Unix Domain Socket vs TCP localhost lab

Pitfalls we hit (or avoided)

  1. “Always set TCP_NODELAY” — wrong for write-heavy tiny-chunk apps without app buffering.
  2. Measuring only connect-per — handshake hides the knob.
  3. Looking only at mean — Nagle+delayed ACK is a tail/mean story.
  4. Assuming WAN 40 ms will appear on localhost — it did not under default quickack here.
  5. Forgetting both ends — server echo path also Nagles unless you set NODELAY on accepted sockets.

Practical checklist

  • Know whether the app issues many small sends or one buffered write.
  • For latency-sensitive reuse RPC: try TCP_NODELAY and measure p50/p95, not only mean.
  • For throughput of tiny writes: prefer app buffering / TCP_CORK over blind NODELAY.
  • Confirm with getsockopt and, if needed, segment counters (/proc/net/snmp OutSegs).
  • Same-host and you care about µs: also consider UDS (separate lab).

Verdict

TCP_NODELAY is not a free latency cheat code. On this box it won chatty 1-byte reuse (~124k vs 73k~~ RPS) and lost batched tiny writes (~~1.7× more segments, ~half the RPS). Under delayed-ACK pressure, Nagle’s mean lied while p50 looked calm. Measure the write pattern you actually ship.

Evidence path on the lab box: lab-evidence/18-tcp-nodelay/results/. Affiliates: 0.

tcp_nodelaynagle algorithmtcp_corkdelayed acktcp latencylocalhost labsresocket options

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5 echo 127.0.0.1:18310 (Nagle) / :18311 (NODELAY). getsockopt on=1 off=0. Reuse 1B: Nagle mean 13.6µs ~73.4k RPS; NODELAY 8.1µs ~124.1k. Multi 32×16B segs arm: Nagle 41µs ~24.4k OutSegsΔ=14202; NODELAY 79µs ~12.6k OutSegsΔ=24296. WWR+QUICKACK=0: Nagle mean 118.8µs p50 23.4; NODELAY mean 22.3µs p50 17.9. Cork control ~38µs. No classic 40ms stall on default loopback quickack. Affiliates: 0. Evidence: lab-evidence/18-tcp-nodelay/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 42

    TCP_QUICKACK Lab: Delayed ACK, Set-Once, and the Reassert Trap

    Hands-on TCP_QUICKACK lab: default QA=1 on loopback; WWR set-once ~18 µs; reassert QA=1 under Nagle p95 ~48 ms, OutSegs ~13×. Real numbers, no Docker.

    30 Sept 2026

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What Nagle and TCP\_NODELAY actually change
  3. Lab topology
  4. Arm 1 — reuse echo (the “RPC” shape)
  5. Arm 2 — many small writes, then one read (coalesce)
  6. Arm 3 — write-write-read + QUICKACK off
  7. How this sits next to keep-alive and UDS
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove