ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 60

  1. Blog
  2. /Observability & SRE

SO_REUSEPORT vs Single Listen: Independent Queues and ListenOverflows

Hands-on SO_REUSEPORT lab: ss shows 4 listen queues; backlog=8 flood 72% vs 62% OK; ~41% fewer ListenOverflows. Real localhost numbers, no Docker.

Aditya Challa·30 September 2026·6 min read

Lab
On this page
  1. Intro — what this post promises
  2. What SO\_REUSEPORT changes
  3. Lab topology
  4. ss: the picture worth saving
  5. Fast echo: RPS basically tied
  6. Affinity: both modes spread work
  7. Slow-accept flood: where reuseport showed up
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict

Intro — what this post promises

SO_REUSEPORT lets multiple processes each bind() + listen() on the same IP:port. Nginx, Envoy, and many multi-worker servers expose it as reuseport. The folklore says it “scales accept.” This lab measures what actually moved on one box.

This is a localhost lab with measured numbers:

  1. What ss shows: one shared listen queue vs N independent queues.
  2. A fair echo microbench: single-listen workers vs SO_REUSEPORT workers.
  3. Accept affinity (which worker PID answers).
  4. A slow-accept flood where backlog pressure creates ListenOverflows.
  5. What this does not claim (WAN, magic RPS multipliers, older-kernel thundering-herd lore alone).

Related links:

  • HTTP Keep-Alive vs Connection: close lab
  • Nginx limit_req rate-limit lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 multi-process TCP servers bound to 127.0.0.1 only. Modes: inherited single listen() FD vs per-worker SO_REUSEPORT. No Docker. No public bind. Affiliates: 0.

Verdict up front: on this kernel, echo RPS was tied (~9.8k) and accept affinity was already fair without reuseport. The clear, measurable win was independent listen queues in ss and fewer ListenOverflows under a tight backlog + slow accept (e.g. +476 → +394 overflows; 62% → 72% client OK on the backlog=8 flood).


What SO_REUSEPORT changes

Without it, N workers typically share one listening socket (inherited FD). With it, each worker owns a listening socket on the same tuple; the kernel steers new SYNs across those sockets.

Related links:

  • man 7 socket — SO_REUSEPORT
  • nginx listen … reuseport

Mental model that matched our lab:

ChoiceWhat it did here
Single listen FDOne ss LISTEN line; one Send-Q (= backlog) shared by all workers
SO_REUSEPORTFour LISTEN lines on the same port; backlog per socket
Fast echo + connect-per clientBottleneck was client/connect — RPS tied
Slow accept + small backlogOverflow counters moved; reuseport dropped fewer

Pair this with the listen-backlog lab: reuseport does not remove the need for accept capacity; it shards the shock absorber.


Lab topology

ThreadPool clients
        │
        ▼
4 worker processes on 127.0.0.1:PORT
   mode=single     → one listen FD, accept() in each worker
   mode=reuseport  → each worker bind+listen with SO_REUSEPORT

Arms:

  1. Fast 64-byte echo, n=8000, concurrency=64.
  2. Affinity probe: server replies with os.getpid(), n=4000.
  3. Slow accept flood: sleep(delay) after accept, small backlog, concurrent connect flood; deltas from /proc/net/netstat ListenOverflows / ListenDrops.

ss: the picture worth saving

With backlog=2 and 4 workers:

Modess -ltn shape
single1 LISTEN line, Send-Q=2
reuseport4 LISTEN lines, Send-Q=2 each

That is the operational tell. If you expected “four queues” but ss shows one line, you are on shared-listen, not reuseport.


Fast echo: RPS basically tied

n=8000, concurrency 64, payload 64 B, 4 workers:

ModeApprox RPSp50p99
single~98245.37 ms16.7 ms
reuseport~98425.25 ms21.2 ms

No meaningful win. The client was opening a fresh TCP connection per op against a cheap handler — we were not backlog-bound. Honest labs say so.


Affinity: both modes spread work

n=4000 probes, reply = worker PID:

ModeUnique workersShare range
single4~23.3%–26.5%
reuseport4~24.2%–25.9%

On this 6.12 kernel, shared accept() was already balanced. Do not sell reuseport as the only path to “fair workers” without measuring your kernel and workload.


Slow-accept flood: where reuseport showed up

Same flood shape, 4 workers, intentional accept delay so the listen queue matters.

backlog=8, delay=25 ms, n=400 connects (200 client threads):

ModeClient OKOK %ListenOverflows ΔListenDrops Δ
single249 / 40062.2%+476+476
reuseport288 / 40072.0%+394+394

backlog=2, delay=40 ms, n=300:

ModeClient OKOK %LO / LD Δ
single163 / 30054.3%+774
reuseport170 / 30056.7%+456 (~41% fewer overflows)

Healthy control — backlog 128, delay 0, n=300: both modes 300/300, overflow delta 0.

Reading: reuseport’s extra queues absorbed more of the same shock before the kernel dropped. Client OK % still tracks accept capacity (the sleep). Raising backlog or speeding accept remains mandatory — reuseport is not a substitute.


Pitfalls we hit (or avoided)

  1. Claiming an RPS win from a connect-heavy echo — we measured a tie and kept it.
  2. Forgetting to look at ss — the multi-LISTEN fingerprint is the config proof.
  3. Ignoring overflow counters — success % alone understates how hard the kernel worked (ListenOverflows).
  4. Treating reuseport as a WAN feature — it is same-host / same-VIP accept steering.
  5. Skipping percentiles — flood survivors had ugly p95; see evidence JSON.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Practical checklist

  • Multi-worker proxy/app: confirm reuseport (or equivalent) if you want per-worker listen queues.
  • Verify with ss -ltnp — expect N LISTEN lines on the port, not one.
  • Under incidents, check ListenOverflows / ListenDrops alongside CPU.
  • Size backlog and accept/worker capacity; reuseport shards the queue, it does not create accept CPUs.
  • Benchmark your handler; do not cite our echo RPS as a universal multiplier.

Verdict

SO_REUSEPORT on this box bought independent listen queues and fewer ListenOverflows under tight backlog + slow accept (72% vs 62% OK on the backlog=8 flood; ~41% fewer overflows on the backlog=2 arm). It did not multiply tiny-echo RPS, and accept affinity was already fair. Use it when you want sharded accept queues; still fix accept capacity and backlog.

Evidence path on the lab box: lab-evidence/16-so-reuseport/results/. Affiliates: 0.

so_reuseportreuseport vs listenlinux accept queuelistenoverflowsnginx reuseportsresocket backlogmulti-worker listen

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5, 4 workers, 127.0.0.1 only. ss: single=1 LISTEN Send-Q; reuseport=4 LISTEN lines. Echo n=8000 c=64: single ~9824 rps vs reuseport ~9842 rps (tied). Affinity n=4000: both ~fair 23–26% shares. Flood backlog=8 delay=25ms n=400: single 249/400 (62.2%) LO+476; reuseport 288/400 (72.0%) LO+394. Flood backlog=2 delay=40ms n=300: LO+774 vs LO+456 (~41% fewer). Healthy backlog=128: both 300/300 LO+0. Affiliates: 0. Evidence: lab-evidence/16-so-reuseport/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 95

    TCP Listen Backlog Lab: somaxconn, ListenOverflows, and Silent Drops

    Hands-on TCP backlog lab: backlog=1 + slow accept → 34/200 OK, ListenOverflows +357; backlog=128 stays 200/200. Real ss + netstat numbers.

    Observability & SRE · 30 Sept 2026

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What SO\_REUSEPORT changes
  3. Lab topology
  4. ss: the picture worth saving
  5. Fast echo: RPS basically tied
  6. Affinity: both modes spread work
  7. Slow-accept flood: where reuseport showed up
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove