ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 95

  1. Blog
  2. /Observability & SRE

TCP Listen Backlog Lab: somaxconn, ListenOverflows, and Silent Drops

Hands-on TCP backlog lab: backlog=1 + slow accept → 34/200 OK, ListenOverflows +357; backlog=128 stays 200/200. Real ss + netstat numbers.

Aditya Challa·30 September 2026·6 min read

Lab
On this page
  1. Intro — what this post promises
  2. What the listen backlog is
  3. Lab topology
  4. Arm A — backlog=1, slow accept: drops show up
  5. Arm B — backlog=128, fast accept: clean sheet
  6. Arm C — backlog=5, medium delay: partial survival
  7. How to diagnose this in production
  8. Pitfalls
  9. Practical checklist
  10. Verdict

Intro — what this post promises

“Connection timed out” on a server that still has free CPU is a classic backlog story. The listen queue filled, the kernel dropped SYNs, and your client blamed the network.

This post is a localhost lab with measured numbers:

  1. What listen(backlog) and somaxconn actually cap.
  2. How to read ListenOverflows / ListenDrops in /proc/net/netstat.
  3. What ss shows for LISTEN Recv-Q / Send-Q.
  4. Three arms: tiny backlog + slow accept, healthy backlog + fast accept, medium backlog + medium delay.
  5. Why clients often see Timeout, not a clean ECONNREFUSED, when tcp_abort_on_overflow=0.

Related links:

  • HTTP Keep-Alive vs Connection: close lab
  • Nginx limit_req rate-limit lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). Python 3.13.5 slow-accept TCP servers bound to 127.0.0.1 only. Concurrent ThreadPool flood clients. Sysctl baseline: somaxconn=4096, tcp_max_syn_backlog=1024, tcp_abort_on_overflow=0, tcp_syncookies=1. No Docker. No public bind. Affiliates: 0. This is an owned lab against our own listener — not an attack playbook.

Verdict up front: backlog is a shock absorber, not a substitute for accept capacity. With listen(1) and an 80 ms accept delay we got 34/200 successes and +357 ListenOverflows. With listen(128) and fast accept, the same flood was 200/200 with zero overflow delta.


What the listen backlog is

When a server calls listen(fd, backlog), the kernel keeps a finite queue of connections waiting to be accept()ed. On Linux the effective cap is also bounded by net.core.somaxconn. Separately, tcp_max_syn_backlog sizes the SYN (incomplete handshake) side.

Related links:

  • man 2 listen
  • kernel doc — ip-sysctl (somaxconn / tcp_max_syn_backlog)

Mental model that matched our lab:

Knob / signalWhat it did here
listen(N)Advertised backlog; ss LISTEN Send-Q showed N (1 / 5 / 128)
somaxconn=4096Upper cap; our N values were all below it
LISTEN Recv-Q (ss)Current queue depth of established-not-yet-accepted conns
ListenOverflows / ListenDropsKernel counters when the accept queue cannot take more
tcp_abort_on_overflow=0Overflow → drop/retry behavior; clients often time out
tcp_syncookies=1SYN cookies can engage under pressure (we saw small SyncookiesSent deltas)

Nginx and other servers expose this as listen … backlog=N;. Same idea: if workers never accept fast enough, raising N only delays the cliff.


Lab topology

ThreadPool flood (n=200, 150 workers)
        │
        ▼
slow_accept_server.py   127.0.0.1:PORT
   listen(backlog)
   sleep(accept_delay_ms) between accepts
   reply "OK\n"

We compared three arms against the same flood shape (200 connects, 150 workers, 2.0 s socket timeout):

ArmBacklogAccept delayIntent
A180 msForce overflow
B1280Healthy control
C540 msMedium pressure

Counters: client ok/fail + deltas from /proc/net/netstat TcpExt + ss -ltn on the listen port.


Arm A — backlog=1, slow accept: drops show up

listen(1), 80 ms sleep between accepts, flood n=200:

MetricValue
Client OK34 / 200 (17%)
Client fail166 (mostly Timeout)
ListenOverflows Δ+357
ListenDrops Δ+357
SyncookiesSent Δ+6
OK latency p50~242 ms

ss showed Send-Q=1 (the backlog) on the LISTEN line. During a repeat flood we sampled Recv-Q=2 mid-burst — the queue was full while accepts crawled.

Honesty: overflow count (357) can exceed client count (200) because the kernel counts dropped SYNs across retransmits/retries, not “one drop per client thread.”


Arm B — backlog=128, fast accept: clean sheet

listen(128), 0 ms accept delay, same flood:

MetricValue
Client OK200 / 200 (100%)
Client fail0
ListenOverflows Δ0
ListenDrops Δ0
OK latency p50~2.0 ms
Wall~0.034 s

ss showed Send-Q=128. Same client pressure, different accept path — no overflow theater. This is the control that proves the flood itself is not “broken networking.”


Arm C — backlog=5, medium delay: partial survival

listen(5), 40 ms accept delay:

MetricValue
Client OK62 / 200 (31%)
Client fail138 (Timeout)
ListenOverflows Δ+303
ListenDrops Δ+303
OK latency p50~284 ms

Bigger backlog than Arm A bought more successes (62 vs 34) but did not eliminate overflows under a slow accept loop. Shock absorber, not a free pass.


How to diagnose this in production

  1. ss -ltn on the service port — LISTEN Recv-Q climbing toward Send-Q (backlog) is the smoking gun.
  2. nstat / /proc/net/netstat — rising ListenOverflows / ListenDrops.
  3. Application accept metrics — accept latency, worker saturation, event-loop blocks.
  4. Client symptom — with tcp_abort_on_overflow=0, expect timeouts and retries more often than immediate refused connections.

Related links:

  • How to read server monitoring graphs
  • Why your average latency graph is lying (p50 / p95 / p99)

Pair with connection reuse: a keep-alive / pool-friendly client lowers connect storms that fill the queue in the first place.

Related links:

  • HTTP Keep-Alive vs Connection: close lab

Pitfalls

  1. Raising somaxconn only — useless if listen() still passes a tiny backlog, or if accept is single-threaded and blocked.
  2. Reading Recv-Q/Send-Q backwards — on LISTEN sockets, Send-Q is the backlog; Recv-Q is current depth.
  3. Assuming overflow ⇒ ECONNREFUSED — not with tcp_abort_on_overflow=0; we mostly saw Timeout.
  4. Ignoring SYN cookies — under pressure, cookie counters can move; do not confuse them with “everything is fine.”
  5. Load-testing someone else’s listener — do not. This lab used owned localhost servers only.

Practical checklist

  • Confirm listen backlog (app config / nginx listen … backlog= / framework server defaults).
  • Confirm somaxconn ≥ that backlog.
  • Watch ss Recv-Q vs Send-Q under load spikes.
  • Alert on rising ListenOverflows/ListenDrops.
  • Fix accept capacity (workers, non-blocking accept loop) before celebrating a larger backlog.
  • Reduce connect churn (pools / keep-alive) so the queue sees fewer storms.

Verdict

On this box, a tiny backlog + slow accept turned a 200-connect localhost flood into 17% success and hundreds of ListenOverflows. The same flood against a backlog-128 fast acceptor was 100% OK with zero overflow delta. Size the queue for bursts — then make sure something actually calls accept() fast enough to drain it.

Evidence path on the lab box: lab-evidence/15-tcp-backlog/results/. Affiliates: 0.

tcp listen backlogsomaxconnlistenoverflowslistendropssyn queuess recv-qsreconnection refused timeout

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5 slow-accept servers on 127.0.0.1 only. Sysctl: somaxconn=4096, tcp_max_syn_backlog=1024, tcp_abort_on_overflow=0, tcp_syncookies=1. Flood n=200 workers=150 timeout=2s. Arm A listen(1)+accept_delay 80ms: ok=34/200 (17%), fail=166 timeouts, ListenOverflows +357, ListenDrops +357, SyncookiesSent +6. Arm B listen(128)+fast accept: ok=200/200, LO/LD +0, p50_ok ~2.0 ms. Arm C listen(5)+40ms delay: ok=62/200 (31%), LO/LD +303. Mid-flood ss Recv-Q sampled at 2 with Send-Q(=backlog) 1. Clients saw Timeout more than ECONNREFUSED when abort_on_overflow=0. Affiliates: 0. Evidence: lab-evidence/15-tcp-backlog/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 60

    SO_REUSEPORT vs Single Listen: Independent Queues and ListenOverflows

    Hands-on SO_REUSEPORT lab: ss shows 4 listen queues; backlog=8 flood 72% vs 62% OK; ~41% fewer ListenOverflows. Real localhost numbers, no Docker.

    Observability & SRE · 30 Sept 2026

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What the listen backlog is
  3. Lab topology
  4. Arm A — backlog=1, slow accept: drops show up
  5. Arm B — backlog=128, fast accept: clean sheet
  6. Arm C — backlog=5, medium delay: partial survival
  7. How to diagnose this in production
  8. Pitfalls
  9. Practical checklist
  10. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove