ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 42

  1. Blog

TCP_QUICKACK Lab: Delayed ACK, Set-Once, and the Reassert Trap

Hands-on TCP_QUICKACK lab: default QA=1 on loopback; WWR set-once ~18 µs; reassert QA=1 under Nagle p95 ~48 ms, OutSegs ~13×. Real numbers, no Docker.

Aditya Challa·30 September 2026·7 min read

Summary
On this page
  1. Intro — what this post promises
  2. What TCP\_QUICKACK actually changes
  3. Lab topology
  4. Arm A — reuse 1-byte echo
  5. Arm B — write-write-read (the delayed-ACK shape)
  6. Arm C — multi tiny writes + OutSegs
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict

Intro — what this post promises

TCP_QUICKACK is the ACK-side twin of the Nagle / TCP_NODELAY story. Folklore says “turn on quickack to kill delayed ACK.” On Linux the option is best-effort, often not sticky, and on this loopback the default already reported 1.

This is a hands-on lab with measured numbers:

  1. What getsockopt(TCP_QUICKACK) returns after connect and after one echo.
  2. Reuse 1-byte echo matrix: Nagle vs TCP_NODELAY × default / set-once quickack.
  3. Write-write-read (WWR): set-once qa0 vs qa1 vs reassert after every send.
  4. Many tiny writes: OutSegs deltas when you thrash quickack under Nagle.
  5. How this sits next to the TCP_NODELAY lab (sender coalesce ≠ ACK policy).

Related links:

  • TCP_NODELAY vs Nagle localhost lab
  • Unix Domain Socket vs TCP localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • HTTP Keep-Alive vs Connection: close lab

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 TCP echo on 127.0.0.1. No Docker. No public bind. Affiliates: 0. This is a loopback microbench, not a WAN delayed-ACK capture.

Verdict up front: default QUICKACK after connect was 1. WWR under Nagle with set-once qa0/qa1 stayed ~18–20 µs p50. Reasserting QUICKACK=1 after every send under Nagle blew mean to ~12 ms (mean/p50 ~326×, p95 ~48 ms) and ~13× more OutSegs on multi-write.


What TCP_QUICKACK actually changes

Delayed ACK lets a receiver wait briefly before acknowledging a single small segment, hoping to piggyback on a reverse data segment. TCP_QUICKACK asks the stack for quick acknowledgements. On Linux it is documented as a mode that the stack may clear — not a permanent “disable delayed ACK forever” switch.

Related links:

  • man 7 tcp — TCP_QUICKACK
  • TCP_NODELAY vs Nagle localhost lab
KnobLayerWhat we measured here
TCP_NODELAYsender coalesce (Nagle)separate lab + reuse matrix here
TCP_QUICKACKreceiver ACK timingset-once vs reassert on loopback
App buffering / TCP_CORKdeliberate batchnot the focus of this post

Lab topology

Echo server + client on 127.0.0.1
Arms: reuse 1 B · WWR (2×32 B then read) · multi 16×8 B
QUICKACK modes: untouched default · set-once 0/1 · reassert 0/1 after each send
Also: getsockopt probe; /proc/net/snmp OutSegs Δ on multi-write

Probe: after create_connection, getsockopt(TCP_QUICKACK) returned 1. Set to 0 → 0. Set to 1 → 1. After one echo it still read 1.


Arm A — reuse 1-byte echo

n=2000 timed RTTs (after warmup):

Armp50p95meanRPS
Nagle + default QA10.3 µs17.714.1~71k
Nagle + qa0 once11.122.915.1~66k
Nagle + qa1 once14.216.316.4~61k
NODELAY + default QA10.7 µs16.812.5~80k
NODELAY + qa0 once11.125.016.4~61k
NODELAY + qa1 once10.719.112.9~78k

On this chatty reuse shape, NODELAY still mattered more than set-once quickack. Default QA already looked “on.”


Arm B — write-write-read (the delayed-ACK shape)

Two 32 B writes, then one read. Classic place Nagle waits on an ACK.

Armp50p95meanmean/p50RPS
Nagle + default QA29.9 µs37.540.31.35~25k
Nagle + qa0 once20.1 µs29.421.21.05~47k
Nagle + qa1 once17.7 µs19.718.81.06~53k
Nagle + qa0 reassert19.833.925.31.28~39k
Nagle + qa1 reassert36.5~48 ms~11.9 ms~326~84
NODELAY + qa0 once16.522.320.61.24~49k
NODELAY + qa1 once26.649.939.01.46~26k

Set-once qa0 vs qa1: both fine; no classic ~40 ms stall on this loopback. The failure mode we actually hit was reasserting QUICKACK=1 after every send under Nagle — p95 jumped into the tens of milliseconds and mean lied hard.

Stability×3 for WWR Nagle qa0 once p50: 18.9 / 31.3 / 17.6 µs (median 18.9).

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Arm C — multi tiny writes + OutSegs

16 writes of 8 B, then drain. n=500.

Armp50meanOutSegs ΔRPS
Nagle + default QA46 µs583886~17k
Nagle + qa0 once50602896~17k
Nagle + qa1 once50642790~16k
Nagle + qa1 reassert~44 ms~37 ms36199 (~13×)~27
NODELAY + qa0 once129 µs14211711~7k

Reassert-qa1 under Nagle did not “fix latency”; it fragmented the session (far more segments) and collapsed RPS. NODELAY still emitted more segments than Nagle set-once — same qualitative story as the NODELAY lab.


How to read these numbers

  • Default QUICKACK=1 on this stack after connect — “enable quickack” is not always a change.
  • Set-once qa0/qa1 barely moved WWR on loopback; do not paste WAN 40 ms folklore as our result.
  • Reassert every write is a different experiment — and it was the pathological arm here.
  • Pair with p95 / mean÷p50, not only p50, before declaring victory.

Pitfalls we hit (or avoided)

  1. Assuming QUICKACK is sticky — man page says the stack may switch modes; reassert behavior is its own hazard.
  2. Expecting a 40 ms localhost stall from set-once qa0 — we did not see it; honesty over folklore.
  3. Thrashing setsockopt(QUICKACK,1) in a hot send loop — measured disaster under Nagle.
  4. Confusing ACK policy with Nagle — fix the write pattern / NODELAY / buffering first.
  5. Reporting only mean — reassert arm’s mean/p50 gap is the teachable failure.

Practical checklist

  • getsockopt(TCP_QUICKACK) after connect before claiming defaults.
  • Prefer app buffering or careful NODELAY over ACK-knob folklore for tiny writes.
  • If you set QUICKACK, prefer set-once (or after known quiet points), not every send.
  • Publish p50 + p95 + mean/p50 for WWR-shaped RPCs.
  • Same-host µs paths: also consider UDS (separate lab).

Related links:

  • Unix Domain Socket vs TCP localhost lab

Verdict

On this loopback box, TCP_QUICKACK defaults to 1 after connect. WWR under Nagle with set-once qa0/qa1 stayed ~18–20 µs p50. Reasserting QUICKACK=1 after every send under Nagle produced mean ~12 ms, p95 ~48 ms, and ~13× OutSegs on multi-write. Treat quickack as a sharp, best-effort tool — not a hot-path toggle.

Evidence path on the lab box: lab-evidence/30-tcp-quickack/results/. Affiliates: 0.

tcp_quickackdelayed acknagle algorithmtcp_nodelaywrite-write-readlocalhost labsrelinux tcp

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5 TCP echo 127.0.0.1. getsockopt default QUICKACK=1 after connect. Reuse 1 B: Nagle default ~71k RPS vs NODELAY default ~80k. WWR Nagle set-once qa0 p50 20.1 µs / qa1 17.7 µs (no classic 40 ms). WWR Nagle qa1 reassert mean 11.9 ms mean/p50 326× p95 ~48 ms RPS ~84. Multi 16×8 B Nagle qa1 reassert OutSegsΔ 36199 vs ~2.8–3.9k set-once; p50 ~44 ms. Stability WWR qa0 p50 median 18.9 µs. No Docker. Affiliates: 0. Evidence: lab-evidence/30-tcp-quickack/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 51

    TCP_NODELAY vs Nagle: Localhost Lab with Real Numbers

    Hands-on TCP_NODELAY vs Nagle: 1-byte reuse 124k vs 73k RPS; 32×16B writes favor Nagle (~1.7× fewer OutSegs). Delayed-ACK mean trap on localhost.

    Observability & SRE · 30 Sept 2026

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 75

    uuid.uuid4 vs uuid.uuid1: Localhost Lab

    Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What TCP\_QUICKACK actually changes
  3. Lab topology
  4. Arm A — reuse 1-byte echo
  5. Arm B — write-write-read (the delayed-ACK shape)
  6. Arm C — multi tiny writes + OutSegs
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove