Plate 66
fork COW RSS vs spawn: Python Multiprocessing Lab
Hands-on fork COW vs multiprocessing spawn lab: PSS/Private_Dirty after fork and page dirtying, plus spawn start latency. Real numbers on the lab box.
Aditya Challa6 min read
On this page
- Intro — what this post promises
- What COW should look like in metrics
- Lab topology
- Arm A — COW memory after fork + half dirty
- Arm B — multiprocessing fork vs spawn start cost
- How to read these numbers
- Pitfalls we hit (or avoided)
- Practical checklist
- Versions / environment pinned
- Why Shared\_Dirty showed \~137 MiB
- Methodology footnote
- Verdict
Intro — what this post promises
os.fork() is “free” until the child writes. Folklore says copy-on-write keeps RSS flat; then people glance at VmRSS and get confused. Separately, multiprocessing spawn re-imports your program — and feels slow next to fork.
This is a hands-on lab with measured numbers:
- Touch a 128 MiB anonymous buffer, fork, snapshot memory.
- Child dirties half the pages; compare VmRSS vs PSS vs Private_Dirty.
- Time
multiprocessingfork vs spawn Process start+join for a tiny worker. - Honest limits on a shared overlay / virtio lab box with tight MemAvailable.
Related links:
- process vs thread pool GIL localhost lab
- context-switch pipe ping-pong localhost lab
- fsync vs fdatasync localhost lab
- SQLite WAL vs DELETE journal localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. MemAvailable ≈ 1.3 GiB at start (box was busy). No Docker. Affiliates: 0. This is not a bare-metal NUMA paper; cgroup accounting can differ on your orchestrator.
Verdict up front: after fork, PSS ~halved (~70 vs ~140 MiB) while VmRSS stayed huge. Dirtying half the buffer raised child Private_Dirty by ~64 MiB; VmRSS moved only +64 kB. spawn start+join was ~16.6× slower than fork (61.4 ms vs 3.70 ms p50).
What COW should look like in metrics
| Metric | After fork (shared pages) | After child writes |
|---|---|---|
| VmRSS | Often still “large” in both | May barely move |
| PSS | ~half of unique charge | Rises as pages go private |
| Shared_Dirty | High (parent+child) | Falls as pages split |
| Private_Dirty | Low in child | Rises ~bytes written |
Related links:
Lab topology
Script: lab-evidence/37-fork-cow-rss-vs-spawn/results/run_lab.py.
Arm A — COW memory after fork + half dirty
Primary run (kB unless noted):
| Snapshot | VmRSS | PSS | Shared_Dirty | Private_Dirty |
|---|---|---|---|---|
| parent after touch | 144948 | 140050 | 0 | 137904 |
| child after fork | 142316 | 70118 | 137108 | 812 |
| child after dirty half | 142380 | 102936 | 71504 | 66416 |
Deltas on the child for half-buffer write:
- Private_Dirty +65604 kB ≈ 64 MiB (matches half of 128 MiB)
- PSS +32818 kB
- VmRSS +64 kB ← the metric that lies if you only watch RSS
Fork latency (parent os.fork call): primary 3.68 ms; repeats 2.09 / 2.82 / 2.70 ms. Child dirty-half loop ~54 ms wall (page fault / copy work).
Related links:
Arm B — multiprocessing fork vs spawn start cost
Empty worker that only reports VmRSS via a Queue:
| Start method | start+join p50 | p95 | child VmRSS p50 |
|---|---|---|---|
| fork | 3.70 ms | 5.10 ms | 12500 kB |
| spawn | 61.4 ms | 82.6 ms | 16206 kB |
spawn / fork ≈ 16.6× on p50 latency. Spawn also showed a fatter fresh interpreter RSS for the toy worker. On macOS/Windows, spawn (or forkserver) is the reality — Linux fork speed is not portable.
How to read these numbers
- VmRSS ≠ unique memory after fork. Use PSS / Private_Dirty (smaps_rollup).
- COW defers copies until write — half dirty ≈ half private growth here.
- fork is fast to start; spawn pays import/bootstrap.
- Container honesty: shared lab pressure (MemAvailable ~1.3 GiB) and overlay/virtio mean absolute RSS baselines include other agents — deltas are the story.
Related links:
- fsync vs fdatasync localhost lab
- nice / ionice CPU and disk priority lab
- ulimit soft vs hard EMFILE lab
Pitfalls we hit (or avoided)
- Declaring “COW failed” because VmRSS stayed flat/high — look at Private_Dirty.
- Forking with threads — not tested; unsafe patterns exist.
- Assuming spawn is “just as fast on Linux” — 16.6× says otherwise here.
- Huge allocations on a memory-tight box — we used 128 MiB because MemAvailable allowed it; watch OOM.
- Comparing to Docker
--ipcfolklore without measuring PSS.
Practical checklist
- After fork-heavy designs, monitor PSS (or cgroup memory) not only VmRSS.
- Avoid parent writing huge shared heaps before forking workers if children will mutate them.
- Prefer forkserver/spawn when you need safety cross-platform — budget ~tens of ms startup here.
- Keep large read-only tables shared; mutate via messages (see pipe/tmpfile lab) when possible.
- Re-measure under your orchestrator’s memory accounting.
Preloading big read-mostly data before fork can be a win until someone writes those pages — the Private_Dirty cliff is the bill.
Related links:
- pipe vs tmpfile IPC localhost lab
- zstd vs gzip vs lz4 compression localhost lab
- SHA-256 vs BLAKE2b vs xxHash localhost lab
Versions / environment pinned
- Python 3.13.5 (
multiprocessingstart methods: fork, spawn, forkserver; default fork) - Buffer 128 MiB
bytearray, 4 KiB page touches - Kernel 6.12 overlay lab VM; MemAvailable ~1.3 GiB at run start
Why Shared_Dirty showed ~137 MiB
Parent had already written every page (buf[off]=1), so pages were anonymous dirty before fork. After fork they became shared dirty between parent and child (~137108 kB Shared_Dirty in the child snapshot). The child’s Private_Dirty stayed tiny (812 kB) until it wrote again.
When the child overwrote the first half, those pages faulted into private copies: Shared_Dirty fell to ~71504 kB, Private_Dirty rose to ~66416 kB. That arithmetic is the COW lecture in three numbers.
Parent VmRSS after the child exited stayed ~145 MiB — the parent still owned its copies; we did not shrink the mapping.
Methodology footnote
COW path uses bytearray so pages are anonymous and mutable (unlike a read-only bytes object). Every page is touched with a single-byte write before fork so we do not measure lazy zero-fill surprises. Child reports JSON over a pipe then _exits; parent waitpids before its own dirty loops. Start-method benches use multiprocessing.get_context(method).Process with a Queue round-trip so we time “usable worker,” not only start() returning.
Workers that only read parent data keep COW shares intact; workers that rewrite large heaps pay Private_Dirty quickly — sometimes enough that spawn + lean imports is healthier than fork + fat parent. Measure PSS under load before picking a start method for a fleet.
Verdict
Forked children shared a touched 128 MiB buffer: PSS ~70 MiB vs parent ~140 MiB with ~137 MiB Shared_Dirty. Writing half the pages cost 64 MiB Private_Dirty~~ while VmRSS barely twitched. Starting a tiny multiprocessing worker took ~3.7 ms with fork vs ~61 ms with spawn (~~16.6×). Trust COW — but trust PSS/Private_Dirty more than RSS headlines.
Evidence path on the lab box: lab-evidence/37-fork-cow-rss-vs-spawn/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5; 128 MiB touched bytearray; MemAvailable ~1.3 GiB on shared box. After fork child PSS ~70 MiB vs parent ~140; Shared_Dirty ~137 MiB. Child dirty half pages: Private_Dirty +65604 kB (~64 MiB), PSS +32818 kB; VmRSS barely moved (+64 kB). fork Process start+join p50 3.70 ms vs spawn 61.4 ms (~16.6x). Honest container/virtio limits. Affiliates: 0. Evidence: lab-evidence/37-fork-cow-rss-vs-spawn/.
Related links
Plate 32
array.array vs list[int] vs bytes Lab
A hands-on Linux localhost lab comparing memory density and numeric throughput for list[int], array.array('i'), bytearray, and memoryview.
30 Sept 2026
Plate 17
platform vs os.uname Inventory: Localhost Lab
Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 75
uuid.uuid4 vs uuid.uuid1: Localhost Lab
Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026