ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 89

  1. Blog
  2. /Observability & SRE

epoll vs select Toy: Ready-Wait Latency and the FD_SETSIZE Cliff

Hands-on epoll vs select toy: ready-wait flat ~3.5 µs (8→900 FDs); select scales 4.6→16.9 µs then hits fd≥1024 cliff. Python numbers, no Docker.

Aditya Challa·30 September 2026·5 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — ready-wait latency (p50 µs)
  5. Arm B — the FD\_SETSIZE-class cliff
  6. When select or poll still show up
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict

Intro — what this post promises

“Just use epoll” is easy advice. Seeing why select falls over — both the O(n) ready-wait cost and the hard FD_SETSIZE / fd≥1024 cliff — sticks better with numbers.

This is a hands-on toy lab with measured numbers:

  1. Ready-wait latency when 1 of N idle sockets becomes readable (epoll / poll / select via Python selectors).
  2. How latency scales from 8 → 900 watched FDs.
  3. Where SelectSelector raises filedescriptor out of range.
  4. That epoll still registered 4096 watched server FDs on this box.

Related links:

  • ulimit nofile / too many open files lab
  • SO_REUSEPORT vs single listen lab
  • TCP listen backlog / somaxconn lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. selectors.DefaultSelector → EpollSelector. Localhost TCP socket pairs only. RLIMIT_NOFILE soft raised to 524288 for the ceiling arms. Toy — not nginx, libuv, or a production event loop. No Docker. Affiliates: 0.

Verdict up front: epoll ready-wait stayed flat at ~3.5 µs from 8→900 FDs. select grew 4.6 → 16.9 µs (8→256) then hit the fd≥1024 cliff. poll scaled without that cliff (3.6 → 32.5 µs).


What we compared

BackendPython classNotes
epollEpollSelectorO(ready) wakeup; no FD_SETSIZE bitmap
pollPollSelectorScales with interest set; no fd-number 1024 cliff here
selectSelectSelectorScans interest set; fails if any fd number ≥ 1024

Related links:

  • man 7 epoll
  • man 2 select

Lab topology

N connected client/server pairs on 127.0.0.1
Watch N server FDs for READ · poke 1 byte on client[0]
200 rounds · report p50 ready-wait latency (µs)
Ceiling: grow N until SelectSelector.select raises

Arm A — ready-wait latency (p50 µs)

Watched FDsepollpollselect
83.43.64.6
643.55.37.2
2563.510.816.9
5123.517.4(skipped — ceiling)
9003.532.5(skipped)

epoll’s line is the lesson: adding idle FDs did not slow the wakeup in this toy. select and poll paid for a larger interest set even when only one socket was hot.


Arm B — the FD_SETSIZE-class cliff

Pairs (watched servers)Max server fdselect()
5001005OK
5121029ValueError: filedescriptor out of range
Forced watch fds ≥1103≥1103same ValueError (register may succeed; select fails)

This is the classic fd number ≥ 1024 failure mode, not “you watched too many interesting sockets” in the abstract. Raising ulimit -n does not raise FD_SETSIZE for select. Epoll registered 4096 watched server FDs successfully on the same box.

Related links:

  • ulimit nofile / too many open files lab

When select or poll still show up

  • Tiny fd sets (a handful of sockets): all three were within a few microseconds — pick the portable API if you must.
  • Older code / POSIX-only builds without epoll: poll avoids the fd-number cliff.
  • One-shot scripts with low connection counts: clarity beats micro-optimizing the waiter.
  • Production servers on Linux: prefer epoll (or an event library that uses it). This lab is the intuition pump, not a framework bake-off.

Connection accept queues and multi-listen designs are separate labs (backlog, SO_REUSEPORT).

Related links:

  • TCP listen backlog / somaxconn lab
  • SO_REUSEPORT vs single listen lab

Pitfalls we hit (or avoided)

  1. Raising ulimit and expecting select to accept fd 2048 — it will not; that is FD_SETSIZE, not RLIMIT_NOFILE.
  2. Counting “connections” but watching only half the FDs — clients still consume fd numbers and can push peers over 1023.
  3. Calling this an nginx RPS result — we measured Python selector ready-wait, not HTTP.
  4. Warming up with leftover bytes — we drained before each poke so latency is wakeup, not accidental data.
  5. Assuming poll is “as flat as epoll” — here poll scaled ~9× from 8→900 FDs.

Practical checklist

  • Linux servers with many connections: use epoll (or a runtime that does).
  • If you still call select, keep all watched fd numbers < 1024 — or stop using select.
  • Pair with ulimit labs: high nofile without epoll still leaves the select cliff.
  • Report watched FD count + max fd when you quote waiter latency.
  • Do not treat toy µs as application RPS.

Verdict

epoll ready-wait stayed ~3.5 µs from 8→900 FDs. select slowed 4.6 → 16.9 µs by 256 FDs, then died when a watched fd reached ≥1024. poll kept working but scaled to 32.5 µs at 900 FDs. Raise ulimit for capacity; switch off select for correctness past the bitmap.

Evidence path on the lab box: lab-evidence/23-epoll-vs-select/results/. Affiliates: 0.

epoll vs selectfd_setsizeepollselectorselect scalabilitylinux event looplocalhost labsrepoll

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5 selectors (DefaultSelector=EpollSelector). Ready-wait p50: epoll flat 3.4–3.5 us from 8→900 watched FDs; poll 3.6→32.5 us; select 4.6→7.2→16.9 us (8/64/256) then SKIP. Select ceiling: n=500 pairs fd max 1005 OK; n=512 fd max 1029 → ValueError filedescriptor out of range; forced fd≥1100 fails select(). Epoll registered 4096 watched FDs OK. Affiliates: 0. Evidence: lab-evidence/23-epoll-vs-select/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — ready-wait latency (p50 µs)
  5. Arm B — the FD\_SETSIZE-class cliff
  6. When select or poll still show up
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove