Plate 89
epoll vs select Toy: Ready-Wait Latency and the FD_SETSIZE Cliff
Hands-on epoll vs select toy: ready-wait flat ~3.5 µs (8→900 FDs); select scales 4.6→16.9 µs then hits fd≥1024 cliff. Python numbers, no Docker.
Aditya Challa5 min read
Intro — what this post promises
“Just use epoll” is easy advice. Seeing why select falls over — both the O(n) ready-wait cost and the hard FD_SETSIZE / fd≥1024 cliff — sticks better with numbers.
This is a hands-on toy lab with measured numbers:
- Ready-wait latency when 1 of N idle sockets becomes readable (
epoll/poll/selectvia Pythonselectors). - How latency scales from 8 → 900 watched FDs.
- Where
SelectSelectorraisesfiledescriptor out of range. - That epoll still registered 4096 watched server FDs on this box.
Related links:
- ulimit nofile / too many open files lab
- SO_REUSEPORT vs single listen lab
- TCP listen backlog / somaxconn lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. selectors.DefaultSelector → EpollSelector. Localhost TCP socket pairs only. RLIMIT_NOFILE soft raised to 524288 for the ceiling arms. Toy — not nginx, libuv, or a production event loop. No Docker. Affiliates: 0.
Verdict up front: epoll ready-wait stayed flat at ~3.5 µs from 8→900 FDs. select grew 4.6 → 16.9 µs (8→256) then hit the fd≥1024 cliff. poll scaled without that cliff (3.6 → 32.5 µs).
What we compared
| Backend | Python class | Notes |
|---|---|---|
| epoll | EpollSelector | O(ready) wakeup; no FD_SETSIZE bitmap |
| poll | PollSelector | Scales with interest set; no fd-number 1024 cliff here |
| select | SelectSelector | Scans interest set; fails if any fd number ≥ 1024 |
Related links:
Lab topology
Arm A — ready-wait latency (p50 µs)
| Watched FDs | epoll | poll | select |
|---|---|---|---|
| 8 | 3.4 | 3.6 | 4.6 |
| 64 | 3.5 | 5.3 | 7.2 |
| 256 | 3.5 | 10.8 | 16.9 |
| 512 | 3.5 | 17.4 | (skipped — ceiling) |
| 900 | 3.5 | 32.5 | (skipped) |
epoll’s line is the lesson: adding idle FDs did not slow the wakeup in this toy. select and poll paid for a larger interest set even when only one socket was hot.
Arm B — the FD_SETSIZE-class cliff
| Pairs (watched servers) | Max server fd | select() |
|---|---|---|
| 500 | 1005 | OK |
| 512 | 1029 | ValueError: filedescriptor out of range |
| Forced watch fds ≥1103 | ≥1103 | same ValueError (register may succeed; select fails) |
This is the classic fd number ≥ 1024 failure mode, not “you watched too many interesting sockets” in the abstract. Raising ulimit -n does not raise FD_SETSIZE for select. Epoll registered 4096 watched server FDs successfully on the same box.
Related links:
When select or poll still show up
- Tiny fd sets (a handful of sockets): all three were within a few microseconds — pick the portable API if you must.
- Older code / POSIX-only builds without epoll: poll avoids the fd-number cliff.
- One-shot scripts with low connection counts: clarity beats micro-optimizing the waiter.
- Production servers on Linux: prefer epoll (or an event library that uses it). This lab is the intuition pump, not a framework bake-off.
Connection accept queues and multi-listen designs are separate labs (backlog, SO_REUSEPORT).
Related links:
Pitfalls we hit (or avoided)
- Raising ulimit and expecting select to accept fd 2048 — it will not; that is FD_SETSIZE, not RLIMIT_NOFILE.
- Counting “connections” but watching only half the FDs — clients still consume fd numbers and can push peers over 1023.
- Calling this an nginx RPS result — we measured Python selector ready-wait, not HTTP.
- Warming up with leftover bytes — we drained before each poke so latency is wakeup, not accidental data.
- Assuming poll is “as flat as epoll” — here poll scaled ~9× from 8→900 FDs.
Practical checklist
- Linux servers with many connections: use epoll (or a runtime that does).
- If you still call
select, keep all watched fd numbers < 1024 — or stop usingselect. - Pair with ulimit labs: high
nofilewithout epoll still leaves the select cliff. - Report watched FD count + max fd when you quote waiter latency.
- Do not treat toy µs as application RPS.
Verdict
epoll ready-wait stayed ~3.5 µs from 8→900 FDs. select slowed 4.6 → 16.9 µs by 256 FDs, then died when a watched fd reached ≥1024. poll kept working but scaled to 32.5 µs at 900 FDs. Raise ulimit for capacity; switch off select for correctness past the bitmap.
Evidence path on the lab box: lab-evidence/23-epoll-vs-select/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5 selectors (DefaultSelector=EpollSelector). Ready-wait p50: epoll flat 3.4–3.5 us from 8→900 watched FDs; poll 3.6→32.5 us; select 4.6→7.2→16.9 us (8/64/256) then SKIP. Select ceiling: n=500 pairs fd max 1005 OK; n=512 fd max 1029 → ValueError filedescriptor out of range; forced fd≥1100 fails select(). Epoll registered 4096 watched FDs OK. Affiliates: 0. Evidence: lab-evidence/23-epoll-vs-select/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026