ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 70

  1. Blog

HTTP Keep-Alive vs Connection: close: Localhost Numbers and a Simulated-RTT Trap

Aditya Challa·30 September 2026·7 min read

Summary
On this page
  1. What keep-alive is (and is not)
  2. Lab topology
  3. Direct localhost: modest RPS gap, loud TIME\_WAIT
  4. Simulated RTT: close costs \~2× wall time
  5. Pitfall: curl-in-a-for-loop is not keep-alive
  6. What we would change in a real service
  7. Reproduce (localhost only)
  8. FAQ

Intro — what this post promises

HTTP/1.1 keep-alive is the boring default that still gets disabled by accident: a client library that sends Connection: close, a proxy that sets keepalive_timeout 0, or a shell script that shells out to curl once per request.

This post is a localhost lab with measured numbers:

  1. What keepalive_timeout actually changes on nginx.
  2. A fair Python http.client comparison: one TCP, many GETs vs new TCP every GET.
  3. Why the localhost RPS gap looks modest — and how a 5 ms + 5 ms delay proxy (simulated RTT) flips it to ~2× wall time.
  4. What concurrent Connection: close floods do to TIME_WAIT.
  5. Why “I timed curl in a for-loop” is not a keep-alive test.

Related links:

  • Nginx limit_req rate-limit lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • OpenTelemetry Collector + Prometheus + Grafana first stack
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). nginx/1.26.3 as a host binary (Debian package). Two nginx masters bound to 127.0.0.1 only on :18110 (keep-alive on) and :18111 (keepalive_timeout 0). A Python delay proxy on :18120 forward to :18110 with delay_connect_ms=5 and delay_req_ms=5 — that is a simulated RTT, not a WAN measurement. Static try_files → 4004-byte index.html. No Docker. No public bind. Affiliates: 0.

Verdict up front: keep-alive is worth keeping on. On pure loopback the RPS win is small; the TIME_WAIT and simulated-RTT arms are where the cost shows up — and where production pain usually starts.


What keep-alive is (and is not)

In HTTP/1.1, persistent connections are the default. The client and server may reuse one TCP socket for many request/response pairs until idle timeout, a hard request cap, an error, or an explicit Connection: close.

Related links:

  • RFC 9112 — HTTP/1.1 (connection management)
  • Module ngx_http_core_module — keepalive_timeout

Mental model that matched our lab:

KnobWhat it did here
keepalive_timeout 65sIdle keep-alive sockets stay reusable; responses advertised Connection: keep-alive
keepalive_requests 10000Cap how many requests one connection may serve before nginx closes
keepalive_timeout 0Server forces Connection: close on every response
Client Connection: closeClient asks to tear down after this response (and pays a new handshake next time)
Shell for i in …; do curl …; doneNew process ⇒ new TCP every time — keep-alive never gets a chance

Keep-alive is not HTTP/2 multiplexing and not a substitute for rate limiting. Pair it with the limit_req lab when abuse is the problem.


Lab topology

Python http.client  ──►  nginx :18110  keepalive_timeout 65s
Python http.client  ──►  nginx :18111  keepalive_timeout 0

Python http.client  ──►  delay proxy :18120  (sleep 5ms connect + 5ms/req)
                              │
                              └──► nginx :18110

Minimal keep-alive server shape:

http {
  keepalive_timeout 65s;
  keepalive_requests 10000;
  server {
    listen 127.0.0.1:18110;
    root /path/to/html;
    location / {
      try_files /index.html =404;
    }
  }
}

Close-forcing twin: same config with keepalive_timeout 0; on :18111.


Direct localhost: modest RPS gap, loud TIME_WAIT

Against the keep-alive server (:18110), Python http.client, n=500 sequential GETs, 4004-byte body:

ModeWallApprox RPSLatency p50
One TCP, reuse0.073 s~68610.107 ms
New TCP each (Connection: close)0.079 s~62930.145 ms

Same client against the close server (:18111):

ModeWallApprox RPS
“Reuse” loop (reconnects when server closes)0.086 s~5785
Explicit close each time0.098 s~5119

Honesty: on loopback, a TCP handshake is cheap. We saw roughly a ~9% RPS drop when the client forced close against the keep-alive server — not a 10× horror story. If your blog post only shows localhost without RTT, you will understate production cost.

What was not modest: socket leftovers. After the keep-alive-server suite, ss TIME_WAIT rows for that port jumped by about +700. After the close-server suite, about +1200. A later concurrent close flood (400 requests, 80 workers) against :18111 left TIME_WAIT around 1603 for that port shortly after the burst. System-wide ss -s reported on the order of ~3910 TIME_WAIT sockets after the floods.

TIME_WAIT is not “wrong” — it is TCP doing its job — but connection-churning clients amplify it. On a busy NAT or a tiny ephemeral-port pool, that is how “random connection failures” show up while CPU still looks fine.


Simulated RTT: close costs ~2× wall time

Loopback hides RTT. We put a tiny Python reverse proxy on :18120 in front of the keep-alive nginx and slept 5 ms on each new TCP accept plus 5 ms before forwarding each request. That is a lab stand-in for “reconnect tax,” not a claim about any cloud region.

Python http.client, n=100 sequential GETs through the delay proxy:

ModeWallApprox RPSLatency p50
One TCP, reuse0.608 s~1645.69 ms
New TCP each (Connection: close)1.211 s~82.611.76 ms

Close / reuse wall ratio: ~1.99×.

That matches the mental model: reuse pays the connect sleep once; close pays it every request, on top of the per-request sleep. Real WANs also pay TLS handshake / congestion — we did not measure TLS in this run.


Pitfall: curl-in-a-for-loop is not keep-alive

We also timed shell loops of the form for i in $(seq 1 N); do curl …; done.

Against direct :18110, 200 curls:

  • default headers: ~1.79 s
  • -H 'Connection: close': ~1.94 s

Through the delay proxy, 50 curls, default vs close were both ~1.05–1.07 s — because each curl process opens its own TCP connection. Keep-alive never engages across process boundaries.

If you want a fair client keep-alive test, use one process with a connection pool (http.client reuse, requests.Session, curl with a single process and multiple URLs carefully, or an HTTP load tool that reuses connections). Measuring process-per-request scripts mostly measures fork + handshake churn.


What we would change in a real service

  1. Leave server keep-alive on unless you have a deliberate reason to force close (some broken middleboxes). On nginx, that is keepalive_timeout greater than zero and a sane keepalive_requests.
  2. Audit clients and proxies for Connection: close defaults — especially sidecar HTTP clients, health-check scripts, and “simple” workers that construct a new client per call.
  3. Watch TIME_WAIT and ephemeral ports under load tests that disable reuse; pair with latency percentiles so a green mean does not hide reconnect spikes.
  4. Do not confuse keep-alive with rate limiting. Abuse controls still need something like limit_req; connection reuse does not stop a determined flood.

Related links:

  • Nginx limit_req rate-limit lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Reproduce (localhost only)

Evidence tree: lab-evidence/12-http-keepalive/ (configs under conf/, clients under scripts/, numbers under results/).

Shape:

  1. Start nginx with temp paths under the lab directory (Debian nginx needs writable client_body_temp_path etc. when not root).
  2. Bind 127.0.0.1 only.
  3. Run scripts/bench.py --port 18110 --label keepalive_server --n 500.
  4. Optionally start scripts/delay_proxy.py and re-bench through :18120.
  5. Snapshot ss -tan state time-wait and ss -s around concurrent close floods.

Do not publish configs that bind 0.0.0.0 without review.


FAQ

Did keep-alive win on localhost?
Yes, but modestly (~6861 vs ~6293 RPS in the n=500 reuse vs close arm). The louder signals were TIME_WAIT growth and the delay-proxy arm.

Is the delay proxy “cheating”?
It is labeled simulated RTT. It exists because pure loopback understates reconnect cost. Treat it as a teaching instrument, not a WAN benchmark.

Did you test HTTP/2 or TLS session tickets?
Not in this lab. Those are separate reuse knobs (multiplexing / session resumption) worth their own measurements.

Any affiliates?
None. No ClickBank, no product CTA.

http-keepaliveconnection-closetime-waitnginxsrelatencyhttp11connection-reuse

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. nginx/1.26.3 host binary, localhost-only :18110 (keepalive_timeout 65s keepalive_requests 10000) and :18111 (keepalive_timeout 0), static 4004-byte try_files. Python http.client: on KA server, sequential reuse n=500 wall 0.073s (~6861 rps, p50 0.107ms) vs Connection: close 0.079s (~6293 rps, p50 0.145ms). Close server TIME_WAIT delta +1200 after suite. Delay proxy :18120→:18110 with delay_connect_ms=5 + delay_req_ms=5: reuse n=100 wall 0.608s (~164 rps, p50 5.69ms) vs close 1.211s (~82.6 rps, p50 11.76ms) ≈1.99× wall. Concurrent close flood 400×80w left TIME_WAIT mid 1603 on :18111. Curl-in-a-shell-loop is not a keep-alive test (new process ⇒ new TCP). ss -s showed timewait ~3910 system-wide after floods. Affiliates: 0. Evidence: lab-evidence/12-http-keepalive/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 30

    Why Your Average Latency Graph Is Lying (p50 / p95 / p99 Playbook)

    Average latency hides tail pain. Learn when mean lies, how p50/p95/p99 work, why averaging quantiles fails, and how Prometheus histograms fix aggregation.

    29 Sept 2026

  • Plate 90

    How to Read Server Monitoring Graphs

    A framework for turning dashboard noise into a diagnosis, based on the four golden signals and the symptom-versus-cause split — summarized from kciter's guide, with the author's hands-on verdict still to come.

    22 Sept 2026

  • How to Read Server Monitoring Graphs

    A framework for diagnosing dashboards: symptoms first, causes second, and why the average response time metric is lying to you.

    10 Sept 2026

On this page

  1. What keep-alive is (and is not)
  2. Lab topology
  3. Direct localhost: modest RPS gap, loud TIME\_WAIT
  4. Simulated RTT: close costs \~2× wall time
  5. Pitfall: curl-in-a-for-loop is not keep-alive
  6. What we would change in a real service
  7. Reproduce (localhost only)
  8. FAQ
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove