Plate 95
TCP Listen Backlog Lab: somaxconn, ListenOverflows, and Silent Drops
Hands-on TCP backlog lab: backlog=1 + slow accept → 34/200 OK, ListenOverflows +357; backlog=128 stays 200/200. Real ss + netstat numbers.
Aditya Challa6 min read
Intro — what this post promises
“Connection timed out” on a server that still has free CPU is a classic backlog story. The listen queue filled, the kernel dropped SYNs, and your client blamed the network.
This post is a localhost lab with measured numbers:
- What
listen(backlog)andsomaxconnactually cap. - How to read ListenOverflows / ListenDrops in
/proc/net/netstat. - What
ssshows for LISTEN Recv-Q / Send-Q. - Three arms: tiny backlog + slow accept, healthy backlog + fast accept, medium backlog + medium delay.
- Why clients often see Timeout, not a clean
ECONNREFUSED, whentcp_abort_on_overflow=0.
Related links:
- HTTP Keep-Alive vs Connection: close lab
- Nginx limit_req rate-limit lab
- Why your average latency graph is lying (p50 / p95 / p99)
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). Python 3.13.5 slow-accept TCP servers bound to 127.0.0.1 only. Concurrent ThreadPool flood clients. Sysctl baseline: somaxconn=4096, tcp_max_syn_backlog=1024, tcp_abort_on_overflow=0, tcp_syncookies=1. No Docker. No public bind. Affiliates: 0. This is an owned lab against our own listener — not an attack playbook.
Verdict up front: backlog is a shock absorber, not a substitute for accept capacity. With listen(1) and an 80 ms accept delay we got 34/200 successes and +357 ListenOverflows. With listen(128) and fast accept, the same flood was 200/200 with zero overflow delta.
What the listen backlog is
When a server calls listen(fd, backlog), the kernel keeps a finite queue of connections waiting to be accept()ed. On Linux the effective cap is also bounded by net.core.somaxconn. Separately, tcp_max_syn_backlog sizes the SYN (incomplete handshake) side.
Related links:
Mental model that matched our lab:
| Knob / signal | What it did here |
|---|---|
listen(N) | Advertised backlog; ss LISTEN Send-Q showed N (1 / 5 / 128) |
somaxconn=4096 | Upper cap; our N values were all below it |
LISTEN Recv-Q (ss) | Current queue depth of established-not-yet-accepted conns |
ListenOverflows / ListenDrops | Kernel counters when the accept queue cannot take more |
tcp_abort_on_overflow=0 | Overflow → drop/retry behavior; clients often time out |
tcp_syncookies=1 | SYN cookies can engage under pressure (we saw small SyncookiesSent deltas) |
Nginx and other servers expose this as listen … backlog=N;. Same idea: if workers never accept fast enough, raising N only delays the cliff.
Lab topology
We compared three arms against the same flood shape (200 connects, 150 workers, 2.0 s socket timeout):
| Arm | Backlog | Accept delay | Intent |
|---|---|---|---|
| A | 1 | 80 ms | Force overflow |
| B | 128 | 0 | Healthy control |
| C | 5 | 40 ms | Medium pressure |
Counters: client ok/fail + deltas from /proc/net/netstat TcpExt + ss -ltn on the listen port.
Arm A — backlog=1, slow accept: drops show up
listen(1), 80 ms sleep between accepts, flood n=200:
| Metric | Value |
|---|---|
| Client OK | 34 / 200 (17%) |
| Client fail | 166 (mostly Timeout) |
ListenOverflows Δ | +357 |
ListenDrops Δ | +357 |
SyncookiesSent Δ | +6 |
| OK latency p50 | ~242 ms |
ss showed Send-Q=1 (the backlog) on the LISTEN line. During a repeat flood we sampled Recv-Q=2 mid-burst — the queue was full while accepts crawled.
Honesty: overflow count (357) can exceed client count (200) because the kernel counts dropped SYNs across retransmits/retries, not “one drop per client thread.”
Arm B — backlog=128, fast accept: clean sheet
listen(128), 0 ms accept delay, same flood:
| Metric | Value |
|---|---|
| Client OK | 200 / 200 (100%) |
| Client fail | 0 |
ListenOverflows Δ | 0 |
ListenDrops Δ | 0 |
| OK latency p50 | ~2.0 ms |
| Wall | ~0.034 s |
ss showed Send-Q=128. Same client pressure, different accept path — no overflow theater. This is the control that proves the flood itself is not “broken networking.”
Arm C — backlog=5, medium delay: partial survival
listen(5), 40 ms accept delay:
| Metric | Value |
|---|---|
| Client OK | 62 / 200 (31%) |
| Client fail | 138 (Timeout) |
ListenOverflows Δ | +303 |
ListenDrops Δ | +303 |
| OK latency p50 | ~284 ms |
Bigger backlog than Arm A bought more successes (62 vs 34) but did not eliminate overflows under a slow accept loop. Shock absorber, not a free pass.
How to diagnose this in production
ss -ltnon the service port — LISTEN Recv-Q climbing toward Send-Q (backlog) is the smoking gun.nstat//proc/net/netstat— rising ListenOverflows / ListenDrops.- Application accept metrics — accept latency, worker saturation, event-loop blocks.
- Client symptom — with
tcp_abort_on_overflow=0, expect timeouts and retries more often than immediate refused connections.
Related links:
Pair with connection reuse: a keep-alive / pool-friendly client lowers connect storms that fill the queue in the first place.
Related links:
Pitfalls
- Raising
somaxconnonly — useless iflisten()still passes a tiny backlog, or if accept is single-threaded and blocked. - Reading Recv-Q/Send-Q backwards — on LISTEN sockets, Send-Q is the backlog; Recv-Q is current depth.
- Assuming overflow ⇒ ECONNREFUSED — not with
tcp_abort_on_overflow=0; we mostly saw Timeout. - Ignoring SYN cookies — under pressure, cookie counters can move; do not confuse them with “everything is fine.”
- Load-testing someone else’s listener — do not. This lab used owned localhost servers only.
Practical checklist
- Confirm
listenbacklog (app config /nginx listen … backlog=/ framework server defaults). - Confirm
somaxconn≥ that backlog. - Watch
ssRecv-Q vs Send-Q under load spikes. - Alert on rising ListenOverflows/ListenDrops.
- Fix accept capacity (workers, non-blocking accept loop) before celebrating a larger backlog.
- Reduce connect churn (pools / keep-alive) so the queue sees fewer storms.
Verdict
On this box, a tiny backlog + slow accept turned a 200-connect localhost flood into 17% success and hundreds of ListenOverflows. The same flood against a backlog-128 fast acceptor was 100% OK with zero overflow delta. Size the queue for bursts — then make sure something actually calls accept() fast enough to drain it.
Evidence path on the lab box: lab-evidence/15-tcp-backlog/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5 slow-accept servers on 127.0.0.1 only. Sysctl: somaxconn=4096, tcp_max_syn_backlog=1024, tcp_abort_on_overflow=0, tcp_syncookies=1. Flood n=200 workers=150 timeout=2s. Arm A listen(1)+accept_delay 80ms: ok=34/200 (17%), fail=166 timeouts, ListenOverflows +357, ListenDrops +357, SyncookiesSent +6. Arm B listen(128)+fast accept: ok=200/200, LO/LD +0, p50_ok ~2.0 ms. Arm C listen(5)+40ms delay: ok=62/200 (31%), LO/LD +303. Mid-flood ss Recv-Q sampled at 2 with Send-Q(=backlog) 1. Clients saw Timeout more than ECONNREFUSED when abort_on_overflow=0. Affiliates: 0. Evidence: lab-evidence/15-tcp-backlog/.
Related links
Plate 60
SO_REUSEPORT vs Single Listen: Independent Queues and ListenOverflows
Hands-on SO_REUSEPORT lab: ss shows 4 listen queues; backlog=8 flood 72% vs 62% OK; ~41% fewer ListenOverflows. Real localhost numbers, no Docker.
Observability & SRE · 30 Sept 2026
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026