ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 77

  1. Blog
  2. /Observability & SRE

ulimit Soft vs Hard File Descriptors: EMFILE Lab with Real Numbers

Hands-on RLIMIT_NOFILE lab: soft=64 → EMFILE after 58 held accepts; soft=512 takes 200/200. Soft raises without root; hard does not. Real numbers.

Aditya Challa·30 September 2026·5 min read

Lab
On this page
  1. Intro — what this post promises
  2. Soft vs hard in one table
  3. Lab topology
  4. Arm A — lower soft, watch EMFILE land on schedule
  5. Arm B — accept server that holds FDs
  6. How to read this next to keep-alive and backlog
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict

Intro — what this post promises

“Too many open files” (EMFILE, errno 24) is the incident that looks like a network outage while CPU is idle. The usual knob people reach for is ulimit -n — and half the time they mix up soft vs hard.

This is a hands-on lab with measured numbers:

  1. What soft and hard RLIMIT_NOFILE mean on a live process.
  2. Hitting EMFILE by opening files and sockets under a lowered soft limit.
  3. Why raising soft within hard needs no root, but raising hard does.
  4. A tiny TCP accept server that holds connections: soft=64 vs soft=512 under the same flood.
  5. What to check in /proc/<pid>/limits during an incident.

Related links:

  • HTTP Keep-Alive vs Connection: close lab
  • Nginx limit_req rate-limit lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). Python 3.13.5 used resource.setrlimit(RLIMIT_NOFILE, …) inside the process only — no host-wide sysctl, no Docker. TCP demo bound to 127.0.0.1 only. Affiliates: 0.

Verdict up front: soft is the ceiling that throws EMFILE. With soft=64 our holding accept server took 58 connections then hit EMFILE; with soft=512 the same 200-client flood was 200/200 with zero EMFILE. Soft raised to 4096 without root; raising hard failed with ValueError: not allowed to raise maximum limit.


Soft vs hard in one table

LimitRole in this lab
SoftEffective ceiling for the process right now
HardCeiling you may raise soft up to without privilege
ulimit -n / -SnSoft
ulimit -HnHard
/proc/<pid>/limits → Max open filesSoft and hard columns for that PID

Related links:

  • man 2 getrlimit / setrlimit
  • man 1 ulimit (bash builtins)

Shell baseline on this box before we touched anything: soft = hard = 524288. That is unusually high for a laptop image and common on some container/CI hosts — which is why teams still ship with soft=1024 in systemd units and get surprised in prod.


Lab topology

Arm A: setrlimit(soft=N) → open files/sockets until OSError errno=24
Arm B: fork server with soft=64 or 512 → hold accepts → client flood n=200

We counted FDs already open via /proc/self/fd so “opened + pre-existing ≈ soft” was visible.


Arm A — lower soft, watch EMFILE land on schedule

SoftWhat we openedCount before EMFILEApprox FDs usedError
256regular files251~257errno 24 Too many open files
64unbound TCP sockets59~65errno 24
1024unbound TCP sockets1019~1025errno 24

Pre-existing FDs in the process were 6 (stdin/out/err + a few). Soft is the budget; the kernel does not care that hard is still 524288.

Permission checks we also ran:

AttemptResult
soft → 4096 (hard unchanged at 524288)OK — no root
soft = hard + 1ValueError: current limit exceeds maximum limit
hard → hard + 1ValueError: not allowed to raise maximum limit

After raising soft to 4096, /proc/self/limits showed: Max open files 4096 524288.


Arm B — accept server that holds FDs

Same client flood: 200 concurrent connects to 127.0.0.1, server keeps accepted sockets open (classic connection leak / FD leak shape).

Server softAccepted & heldEMFILE on further accept?Client OK
6458Yes (first EMFILE after 58 accepts)partial (server stopped taking more)
512200No (0 samples)200 / 200

This is the production story: the listen socket is still up, CPU is fine, and accept() starts returning EMFILE. Clients see stalls, resets, or refusals depending on timing — dashboards scream “network” while ls /proc/<pid>/fd | wc -l is sitting on the soft ceiling.


How to read this next to keep-alive and backlog

Keep-alive reduces FD churn by reusing connections. A connection leak increases FD use until soft kills you. Listen backlog absorbs SYN/accept delay — it does not raise RLIMIT_NOFILE. Three different knobs; three different graphs.

Related links:

  • HTTP Keep-Alive vs Connection: close lab
  • How to read server monitoring graphs

Pitfalls we hit (or avoided)

  1. Raising soft in a shell and expecting systemd services to inherit it — unit LimitNOFILE= is what matters for daemons.
  2. Looking only at hard — hard can be huge while soft=1024 still EMFILEs.
  3. Forgetting already-open FDs — soft=64 does not mean 64 new sockets; we had 6 open first.
  4. Assuming containers match the host ulimit — measure /proc/<pid>/limits inside the workload.
  5. Calling this a cgroup nofile deep dive — we measured process rlimits only.

Practical checklist

  • During “too many open files”: cat /proc/<pid>/limits and ls /proc/<pid>/fd | wc -l.
  • Confirm soft vs hard; raise soft up to hard in the unit/image that runs the process.
  • Fix leaks (held accepts, forgotten close, unbounded connection pools) — a higher ulimit only buys time.
  • For proxies: prefer keep-alive / pools so you need fewer FDs per RPS.
  • Do not confuse EMFILE (process limit) with ENFILE (system-wide file table).

Verdict

Soft RLIMIT_NOFILE is the limit that produces EMFILE. We reproduced it on schedule (soft 256 → 251 files; soft 64 → 59 sockets) and in an accept-hold flood (58 held then EMFILE vs 200/200 at soft 512). Raise soft within hard without root; treat hard raises and leak fixes as separate workstreams.

Evidence path on the lab box: lab-evidence/17-ulimit-fds/results/. Affiliates: 0.

ulimit soft vs hardemfile too many open filesrlimit_nofilefile descriptor limitmax open filessrelinux ulimitaccept emfile

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5. Shell baseline soft=hard=524288. setrlimit: soft=256 files opened 251 then errno 24 EMFILE (ceiling≈257); soft=64 sockets opened 59 then EMFILE; soft=1024 sockets opened 1019 then EMFILE. soft>hard → ValueError; raise soft→4096 OK; raise hard → ValueError not allowed. Holding-accept server flood n=200: soft=64 accepted_and_held=58 then EMFILE; soft=512 accepted 200/200 emfile=0. Affiliates: 0. Evidence: lab-evidence/17-ulimit-fds/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Soft vs hard in one table
  3. Lab topology
  4. Arm A — lower soft, watch EMFILE land on schedule
  5. Arm B — accept server that holds FDs
  6. How to read this next to keep-alive and backlog
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove