Plate 77
ulimit Soft vs Hard File Descriptors: EMFILE Lab with Real Numbers
Hands-on RLIMIT_NOFILE lab: soft=64 → EMFILE after 58 held accepts; soft=512 takes 200/200. Soft raises without root; hard does not. Real numbers.
Aditya Challa5 min read
Intro — what this post promises
“Too many open files” (EMFILE, errno 24) is the incident that looks like a network outage while CPU is idle. The usual knob people reach for is ulimit -n — and half the time they mix up soft vs hard.
This is a hands-on lab with measured numbers:
- What soft and hard
RLIMIT_NOFILEmean on a live process. - Hitting EMFILE by opening files and sockets under a lowered soft limit.
- Why raising soft within hard needs no root, but raising hard does.
- A tiny TCP accept server that holds connections: soft=64 vs soft=512 under the same flood.
- What to check in
/proc/<pid>/limitsduring an incident.
Related links:
- HTTP Keep-Alive vs Connection: close lab
- Nginx limit_req rate-limit lab
- Why your average latency graph is lying (p50 / p95 / p99)
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). Python 3.13.5 used resource.setrlimit(RLIMIT_NOFILE, …) inside the process only — no host-wide sysctl, no Docker. TCP demo bound to 127.0.0.1 only. Affiliates: 0.
Verdict up front: soft is the ceiling that throws EMFILE. With soft=64 our holding accept server took 58 connections then hit EMFILE; with soft=512 the same 200-client flood was 200/200 with zero EMFILE. Soft raised to 4096 without root; raising hard failed with ValueError: not allowed to raise maximum limit.
Soft vs hard in one table
| Limit | Role in this lab |
|---|---|
| Soft | Effective ceiling for the process right now |
| Hard | Ceiling you may raise soft up to without privilege |
ulimit -n / -Sn | Soft |
ulimit -Hn | Hard |
/proc/<pid>/limits → Max open files | Soft and hard columns for that PID |
Related links:
Shell baseline on this box before we touched anything: soft = hard = 524288. That is unusually high for a laptop image and common on some container/CI hosts — which is why teams still ship with soft=1024 in systemd units and get surprised in prod.
Lab topology
We counted FDs already open via /proc/self/fd so “opened + pre-existing ≈ soft” was visible.
Arm A — lower soft, watch EMFILE land on schedule
| Soft | What we opened | Count before EMFILE | Approx FDs used | Error |
|---|---|---|---|---|
| 256 | regular files | 251 | ~257 | errno 24 Too many open files |
| 64 | unbound TCP sockets | 59 | ~65 | errno 24 |
| 1024 | unbound TCP sockets | 1019 | ~1025 | errno 24 |
Pre-existing FDs in the process were 6 (stdin/out/err + a few). Soft is the budget; the kernel does not care that hard is still 524288.
Permission checks we also ran:
| Attempt | Result |
|---|---|
| soft → 4096 (hard unchanged at 524288) | OK — no root |
| soft = hard + 1 | ValueError: current limit exceeds maximum limit |
| hard → hard + 1 | ValueError: not allowed to raise maximum limit |
After raising soft to 4096, /proc/self/limits showed: Max open files 4096 524288.
Arm B — accept server that holds FDs
Same client flood: 200 concurrent connects to 127.0.0.1, server keeps accepted sockets open (classic connection leak / FD leak shape).
| Server soft | Accepted & held | EMFILE on further accept? | Client OK |
|---|---|---|---|
| 64 | 58 | Yes (first EMFILE after 58 accepts) | partial (server stopped taking more) |
| 512 | 200 | No (0 samples) | 200 / 200 |
This is the production story: the listen socket is still up, CPU is fine, and accept() starts returning EMFILE. Clients see stalls, resets, or refusals depending on timing — dashboards scream “network” while ls /proc/<pid>/fd | wc -l is sitting on the soft ceiling.
How to read this next to keep-alive and backlog
Keep-alive reduces FD churn by reusing connections. A connection leak increases FD use until soft kills you. Listen backlog absorbs SYN/accept delay — it does not raise RLIMIT_NOFILE. Three different knobs; three different graphs.
Related links:
Pitfalls we hit (or avoided)
- Raising soft in a shell and expecting systemd services to inherit it — unit
LimitNOFILE=is what matters for daemons. - Looking only at hard — hard can be huge while soft=1024 still EMFILEs.
- Forgetting already-open FDs — soft=64 does not mean 64 new sockets; we had 6 open first.
- Assuming containers match the host ulimit — measure
/proc/<pid>/limitsinside the workload. - Calling this a cgroup
nofiledeep dive — we measured process rlimits only.
Practical checklist
- During “too many open files”:
cat /proc/<pid>/limitsandls /proc/<pid>/fd | wc -l. - Confirm soft vs hard; raise soft up to hard in the unit/image that runs the process.
- Fix leaks (held accepts, forgotten
close, unbounded connection pools) — a higher ulimit only buys time. - For proxies: prefer keep-alive / pools so you need fewer FDs per RPS.
- Do not confuse EMFILE (process limit) with ENFILE (system-wide file table).
Verdict
Soft RLIMIT_NOFILE is the limit that produces EMFILE. We reproduced it on schedule (soft 256 → 251 files; soft 64 → 59 sockets) and in an accept-hold flood (58 held then EMFILE vs 200/200 at soft 512). Raise soft within hard without root; treat hard raises and leak fixes as separate workstreams.
Evidence path on the lab box: lab-evidence/17-ulimit-fds/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5. Shell baseline soft=hard=524288. setrlimit: soft=256 files opened 251 then errno 24 EMFILE (ceiling≈257); soft=64 sockets opened 59 then EMFILE; soft=1024 sockets opened 1019 then EMFILE. soft>hard → ValueError; raise soft→4096 OK; raise hard → ValueError not allowed. Holding-accept server flood n=200: soft=64 accepted_and_held=58 then EMFILE; soft=512 accepted 200/200 emfile=0. Affiliates: 0. Evidence: lab-evidence/17-ulimit-fds/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026