Plate 38
nice / ionice Lab: CPU Nice 19 Slowdown and Idle-Class Disk
Hands-on nice/ionice lab: victim nice19 under 6 hogs ~5.7× slower (p95 ~7.4s); ionice idle O_DIRECT read ~6.1× under writers. Real numbers, no Docker.
Aditya Challa6 min read
Intro — what this post promises
nice and ionice are the knobs people reach for when a batch job should “get out of the way.” Folklore says nice 19 makes you harmless and idle-class IO never steals the disk. This lab measures both.
This is a hands-on lab with measured numbers:
- Fixed-work CPU victim under 6 busy-loop hogs at different nice values.
- Proof that
nice -5did nothing here withoutCAP_SYS_NICE. - Why warm page-cache scans hide
ionice. - O_DIRECT reads/writes under competing best-effort writers: idle vs BE0/BE7.
- Tails: p95 under nice 19 is where the pain lives.
Related links:
- ProcessPool vs ThreadPool GIL localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
- mmap vs read (+ O_DIRECT) localhost lab
- epoll vs select FD_SETSIZE localhost lab
Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 + nice / ionice / dd iflag=direct. No Docker. No GPU. No API keys. drop_caches denied (no root). Affiliates: 0.
Verdict up front: victim nice 19 under 6 nice-0 hogs was 5.7× slower p50 (398 ms → 2256 ms) with p95 ~7.4 s. Hogs at nice 19 left the nice-0 victim near baseline (1.07×). Warm cached reads hid ionice; O_DIRECT under 2 BE0 writers made idle-class reads 6.1× slower vs 2.2× for BE0.
What nice and ionice change
| Tool | Resource | Classes / range we used |
|---|---|---|
nice | CPU scheduling weight | 0 … 19 (user); negative needs privilege |
ionice | Block IO priority (cfq/bfq-era classes still exposed) | best-effort prio 0–7; idle class 3 |
Related links:
Lab topology
Probes: nice -n0/10/19 → reported priority 0/10/19. nice -n-5 → still 0 (no capability). ionice -c3 → idle; -c2 -n0/7 → best-effort: prio 0/7.
Arm A — CPU victim vs nice hogs
Fixed work, n=25 timed runs. Baseline (no hogs, nice 0): p50 ≈ 398 ms, p95 ≈ 627 ms.
| Hogs nice | Victim nice | p50 | p95 | slowdown× p50 |
|---|---|---|---|---|
| (none) | 0 | 398 ms | 627 | 1.00 |
| 0 | 0 | 436 | 722 | 1.09 |
| 0 | 19 | 2256 ms | 7410 ms | 5.67 |
| 19 | 0 | 425 ms | 508 | 1.07 |
Background the batch job (nice 19 on the hogs) and the interactive victim barely noticed. Nice the victim to 19 under aggressive hogs and you pay ~5.7× at p50 — and ~7.4 s p95 on the same fixed work.
Negative nice arms were attempted; without CAP_SYS_NICE the process stayed at priority 0. Do not cite those as “higher priority wins.”
Related links:
Arm B — warm cache: ionice looks fake
128 MiB sequential read already in page cache, 4 competing BE readers:
| Victim ionice | p50 | GB/s | vs warm solo |
|---|---|---|---|
| warm solo BE4 | 28 ms | 4.44 | 1.00 |
| idle (class 3) | 29 ms | 4.38 | 1.01 |
| BE7 | 47 ms | 2.68 | 1.66 |
| BE0 | 36 ms | 3.49 | 1.27 |
Idle ≈ solo. That is not “ionice broken” — it is memory bandwidth, not disk elevator. drop_caches was unavailable without root, so we moved to O_DIRECT.
Related links:
Arm C — O_DIRECT under competing writers
64 MiB dd iflag=direct while 2 BE0 oflag=direct writers hammer separate files.
Reads
| Victim | p50 | MB/s | slowdown× vs solo BE0 |
|---|---|---|---|
| solo BE0 | 51 ms | ~1246 | 1.00 |
| vs2 writers, BE0 | 112 ms | ~572 | 2.18 |
| vs2 writers, BE7 | 111 ms | ~578 | 2.15 |
| vs2 writers, idle | 312 ms | ~205 | 6.08 |
Writes (64 MiB oflag=direct + fsync-ish conv=fsync)
| Victim | p50 | MB/s | slowdown× vs write solo BE0 |
|---|---|---|---|
| solo BE0 | 88 ms | ~728 | 1.00 |
| vs2, BE0 | 153 ms | ~418 | 1.74 |
| vs2, BE7 | 119 ms | ~536 | 1.36 |
| vs2, idle | 339 ms | ~189 | 3.86 |
Idle-class did what it says on the tin once IO actually hit the device: the victim yielded hard. Best-effort stayed in the fight (~2× read slowdown vs ~6× idle).
How to read these numbers
- nice 19 on the batch job protects latency-sensitive work better than hoping the interactive task “wins” by default under equal nice.
- nice 19 on the latency path under load is a self-inflicted p95 tax.
- ionice idle needs real block IO (direct/cold) to show up; page cache lies.
- BE prio 0 vs 7 was close here under this writer pair; idle vs BE was the cliff.
Pitfalls we hit (or avoided)
- Judging ionice on warm buffered reads — looked like a no-op until O_DIRECT.
- Assuming
nice -5works for unprivileged users — probe stayed 0. - Looking only at p50 — nice19 victim p95 was multi-second.
- Nice-ing the wrong process — background the batch hogs, not the request path.
- Claiming elevator theory for a specific IO scheduler — we report class behavior on this box, not a BFQ paper.
Practical checklist
- Batch / backup / scrape:
nice -n19(and considerionice -c3for disk) on the batch side. - Latency path: keep nice near 0; measure p95 under contention.
- Validate ionice with direct or cold IO, not only cached
cat. - Confirm with
nice/ioniceprobes (no silent privilege fail). - Keep evidence next to any “we niced it” capacity claim.
Verdict
On this box, a nice-0 victim under 6 nice-0 hogs was fine (1.1×); the same victim at nice 19 paid 5.7× p50 and ~7.4 s p95. Nicing the hogs to 19 protected the victim (1.07×). Warm cache hid ionice; under O_DIRECT writers, idle-class reads were 6.1× slower vs 2.2× for best-effort. Priority knobs work — aim them at the batch job, and measure the device path.
Evidence path on the lab box: lab-evidence/31-nice-ionice/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. CPU fixed-work victim: baseline nice0 p50 398 ms; 6 hogs nice0 + victim nice19 p50 2256 ms (~5.67×), p95 7410 ms; hogs nice19 + victim nice0 ~1.07×. nice -5 ineffective without CAP_SYS_NICE (probe stayed 0). Warm page-cache scans hid ionice; O_DIRECT 64 MiB: solo BE0 ~1246 MB/s; vs 2 BE0 writers victim idle p50 312 ms (~6.08×) vs victim BE0 112 ms (~2.18×). Write vs2: idle ~3.86× vs BE0 ~1.74×. No Docker. Affiliates: 0. Evidence: lab-evidence/31-nice-ionice/.
Related links
Plate 57
Context-Switch Microbench: Pipe Ping-Pong Process vs Thread
Hands-on ctx-switch lab: process pipe RTT p50 3.2 us (~0.50M sw/s); thread pipe 2.85 us; Event 9.3 us; 6 hogs blow p95 to 19 us. Real numbers, no Docker.
Observability & SRE · 30 Sept 2026
Plate 17
platform vs os.uname Inventory: Localhost Lab
Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 75
uuid.uuid4 vs uuid.uuid1: Localhost Lab
Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026