Plate 70
HTTP Keep-Alive vs Connection: close: Localhost Numbers and a Simulated-RTT Trap
Aditya Challa7 min read
Intro — what this post promises
HTTP/1.1 keep-alive is the boring default that still gets disabled by accident: a client library that sends Connection: close, a proxy that sets keepalive_timeout 0, or a shell script that shells out to curl once per request.
This post is a localhost lab with measured numbers:
- What
keepalive_timeoutactually changes on nginx. - A fair Python
http.clientcomparison: one TCP, many GETs vs new TCP every GET. - Why the localhost RPS gap looks modest — and how a 5 ms + 5 ms delay proxy (simulated RTT) flips it to ~2× wall time.
- What concurrent
Connection: closefloods do to TIME_WAIT. - Why “I timed curl in a for-loop” is not a keep-alive test.
Related links:
- Nginx limit_req rate-limit lab
- Why your average latency graph is lying (p50 / p95 / p99)
- OpenTelemetry Collector + Prometheus + Grafana first stack
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). nginx/1.26.3 as a host binary (Debian package). Two nginx masters bound to 127.0.0.1 only on :18110 (keep-alive on) and :18111 (keepalive_timeout 0). A Python delay proxy on :18120 forward to :18110 with delay_connect_ms=5 and delay_req_ms=5 — that is a simulated RTT, not a WAN measurement. Static try_files → 4004-byte index.html. No Docker. No public bind. Affiliates: 0.
Verdict up front: keep-alive is worth keeping on. On pure loopback the RPS win is small; the TIME_WAIT and simulated-RTT arms are where the cost shows up — and where production pain usually starts.
What keep-alive is (and is not)
In HTTP/1.1, persistent connections are the default. The client and server may reuse one TCP socket for many request/response pairs until idle timeout, a hard request cap, an error, or an explicit Connection: close.
Related links:
Mental model that matched our lab:
| Knob | What it did here |
|---|---|
keepalive_timeout 65s | Idle keep-alive sockets stay reusable; responses advertised Connection: keep-alive |
keepalive_requests 10000 | Cap how many requests one connection may serve before nginx closes |
keepalive_timeout 0 | Server forces Connection: close on every response |
Client Connection: close | Client asks to tear down after this response (and pays a new handshake next time) |
Shell for i in …; do curl …; done | New process ⇒ new TCP every time — keep-alive never gets a chance |
Keep-alive is not HTTP/2 multiplexing and not a substitute for rate limiting. Pair it with the limit_req lab when abuse is the problem.
Lab topology
Minimal keep-alive server shape:
Close-forcing twin: same config with keepalive_timeout 0; on :18111.
Direct localhost: modest RPS gap, loud TIME_WAIT
Against the keep-alive server (:18110), Python http.client, n=500 sequential GETs, 4004-byte body:
| Mode | Wall | Approx RPS | Latency p50 |
|---|---|---|---|
| One TCP, reuse | 0.073 s | ~6861 | 0.107 ms |
New TCP each (Connection: close) | 0.079 s | ~6293 | 0.145 ms |
Same client against the close server (:18111):
| Mode | Wall | Approx RPS |
|---|---|---|
| “Reuse” loop (reconnects when server closes) | 0.086 s | ~5785 |
| Explicit close each time | 0.098 s | ~5119 |
Honesty: on loopback, a TCP handshake is cheap. We saw roughly a ~9% RPS drop when the client forced close against the keep-alive server — not a 10× horror story. If your blog post only shows localhost without RTT, you will understate production cost.
What was not modest: socket leftovers. After the keep-alive-server suite, ss TIME_WAIT rows for that port jumped by about +700. After the close-server suite, about +1200. A later concurrent close flood (400 requests, 80 workers) against :18111 left TIME_WAIT around 1603 for that port shortly after the burst. System-wide ss -s reported on the order of ~3910 TIME_WAIT sockets after the floods.
TIME_WAIT is not “wrong” — it is TCP doing its job — but connection-churning clients amplify it. On a busy NAT or a tiny ephemeral-port pool, that is how “random connection failures” show up while CPU still looks fine.
Simulated RTT: close costs ~2× wall time
Loopback hides RTT. We put a tiny Python reverse proxy on :18120 in front of the keep-alive nginx and slept 5 ms on each new TCP accept plus 5 ms before forwarding each request. That is a lab stand-in for “reconnect tax,” not a claim about any cloud region.
Python http.client, n=100 sequential GETs through the delay proxy:
| Mode | Wall | Approx RPS | Latency p50 |
|---|---|---|---|
| One TCP, reuse | 0.608 s | ~164 | 5.69 ms |
New TCP each (Connection: close) | 1.211 s | ~82.6 | 11.76 ms |
Close / reuse wall ratio: ~1.99×.
That matches the mental model: reuse pays the connect sleep once; close pays it every request, on top of the per-request sleep. Real WANs also pay TLS handshake / congestion — we did not measure TLS in this run.
Pitfall: curl-in-a-for-loop is not keep-alive
We also timed shell loops of the form for i in $(seq 1 N); do curl …; done.
Against direct :18110, 200 curls:
- default headers: ~1.79 s
-H 'Connection: close': ~1.94 s
Through the delay proxy, 50 curls, default vs close were both ~1.05–1.07 s — because each curl process opens its own TCP connection. Keep-alive never engages across process boundaries.
If you want a fair client keep-alive test, use one process with a connection pool (http.client reuse, requests.Session, curl with a single process and multiple URLs carefully, or an HTTP load tool that reuses connections). Measuring process-per-request scripts mostly measures fork + handshake churn.
What we would change in a real service
- Leave server keep-alive on unless you have a deliberate reason to force close (some broken middleboxes). On nginx, that is
keepalive_timeoutgreater than zero and a sanekeepalive_requests. - Audit clients and proxies for
Connection: closedefaults — especially sidecar HTTP clients, health-check scripts, and “simple” workers that construct a new client per call. - Watch TIME_WAIT and ephemeral ports under load tests that disable reuse; pair with latency percentiles so a green mean does not hide reconnect spikes.
- Do not confuse keep-alive with rate limiting. Abuse controls still need something like
limit_req; connection reuse does not stop a determined flood.
Related links:
Reproduce (localhost only)
Evidence tree: lab-evidence/12-http-keepalive/ (configs under conf/, clients under scripts/, numbers under results/).
Shape:
- Start nginx with temp paths under the lab directory (Debian nginx needs writable
client_body_temp_pathetc. when not root). - Bind 127.0.0.1 only.
- Run
scripts/bench.py --port 18110 --label keepalive_server --n 500. - Optionally start
scripts/delay_proxy.pyand re-bench through:18120. - Snapshot
ss -tan state time-waitandss -saround concurrent close floods.
Do not publish configs that bind 0.0.0.0 without review.
FAQ
Did keep-alive win on localhost?
Yes, but modestly (~6861 vs ~6293 RPS in the n=500 reuse vs close arm). The louder signals were TIME_WAIT growth and the delay-proxy arm.
Is the delay proxy “cheating”?
It is labeled simulated RTT. It exists because pure loopback understates reconnect cost. Treat it as a teaching instrument, not a WAN benchmark.
Did you test HTTP/2 or TLS session tickets?
Not in this lab. Those are separate reuse knobs (multiplexing / session resumption) worth their own measurements.
Any affiliates?
None. No ClickBank, no product CTA.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. nginx/1.26.3 host binary, localhost-only :18110 (keepalive_timeout 65s keepalive_requests 10000) and :18111 (keepalive_timeout 0), static 4004-byte try_files. Python http.client: on KA server, sequential reuse n=500 wall 0.073s (~6861 rps, p50 0.107ms) vs Connection: close 0.079s (~6293 rps, p50 0.145ms). Close server TIME_WAIT delta +1200 after suite. Delay proxy :18120→:18110 with delay_connect_ms=5 + delay_req_ms=5: reuse n=100 wall 0.608s (~164 rps, p50 5.69ms) vs close 1.211s (~82.6 rps, p50 11.76ms) ≈1.99× wall. Concurrent close flood 400×80w left TIME_WAIT mid 1603 on :18111. Curl-in-a-shell-loop is not a keep-alive test (new process ⇒ new TCP). ss -s showed timewait ~3910 system-wide after floods. Affiliates: 0. Evidence: lab-evidence/12-http-keepalive/.
Related links
Plate 30
Why Your Average Latency Graph Is Lying (p50 / p95 / p99 Playbook)
Average latency hides tail pain. Learn when mean lies, how p50/p95/p99 work, why averaging quantiles fails, and how Prometheus histograms fix aggregation.
29 Sept 2026
Plate 90
How to Read Server Monitoring Graphs
A framework for turning dashboard noise into a diagnosis, based on the four golden signals and the symptom-versus-cause split — summarized from kciter's guide, with the author's hands-on verdict still to come.
22 Sept 2026
How to Read Server Monitoring Graphs
A framework for diagnosing dashboards: symptoms first, causes second, and why the average response time metric is lying to you.
10 Sept 2026