ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 66

  1. Blog

fork COW RSS vs spawn: Python Multiprocessing Lab

Hands-on fork COW vs multiprocessing spawn lab: PSS/Private_Dirty after fork and page dirtying, plus spawn start latency. Real numbers on the lab box.

Aditya Challa·30 September 2026·6 min read

Summary
On this page
  1. Intro — what this post promises
  2. What COW should look like in metrics
  3. Lab topology
  4. Arm A — COW memory after fork + half dirty
  5. Arm B — multiprocessing fork vs spawn start cost
  6. How to read these numbers
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Versions / environment pinned
  10. Why Shared\_Dirty showed \~137 MiB
  11. Methodology footnote
  12. Verdict

Intro — what this post promises

os.fork() is “free” until the child writes. Folklore says copy-on-write keeps RSS flat; then people glance at VmRSS and get confused. Separately, multiprocessing spawn re-imports your program — and feels slow next to fork.

This is a hands-on lab with measured numbers:

  1. Touch a 128 MiB anonymous buffer, fork, snapshot memory.
  2. Child dirties half the pages; compare VmRSS vs PSS vs Private_Dirty.
  3. Time multiprocessing fork vs spawn Process start+join for a tiny worker.
  4. Honest limits on a shared overlay / virtio lab box with tight MemAvailable.

Related links:

  • process vs thread pool GIL localhost lab
  • context-switch pipe ping-pong localhost lab
  • fsync vs fdatasync localhost lab
  • SQLite WAL vs DELETE journal localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. MemAvailable ≈ 1.3 GiB at start (box was busy). No Docker. Affiliates: 0. This is not a bare-metal NUMA paper; cgroup accounting can differ on your orchestrator.

Verdict up front: after fork, PSS ~halved (~70 vs ~140 MiB) while VmRSS stayed huge. Dirtying half the buffer raised child Private_Dirty by ~64 MiB; VmRSS moved only +64 kB. spawn start+join was ~16.6× slower than fork (61.4 ms vs 3.70 ms p50).


What COW should look like in metrics

MetricAfter fork (shared pages)After child writes
VmRSSOften still “large” in bothMay barely move
PSS~half of unique chargeRises as pages go private
Shared_DirtyHigh (parent+child)Falls as pages split
Private_DirtyLow in childRises ~bytes written

Related links:

  • process vs thread pool GIL localhost lab

Lab topology

bytearray(128 MiB) → touch every page → os.fork()
Child: snapshot → dirty first half of pages → snapshot → report via pipe
Parent: wait → dirty second half → dirty first half
Also: multiprocessing get_context('fork'|'spawn') Process start+join × N
Sources: /proc/self/status VmRSS; /proc/self/smaps_rollup PSS/Private_Dirty/Shared_*

Script: lab-evidence/37-fork-cow-rss-vs-spawn/results/run_lab.py.


Arm A — COW memory after fork + half dirty

Primary run (kB unless noted):

SnapshotVmRSSPSSShared_DirtyPrivate_Dirty
parent after touch1449481400500137904
child after fork14231670118137108812
child after dirty half1423801029367150466416

Deltas on the child for half-buffer write:

  • Private_Dirty +65604 kB ≈ 64 MiB (matches half of 128 MiB)
  • PSS +32818 kB
  • VmRSS +64 kB ← the metric that lies if you only watch RSS

Fork latency (parent os.fork call): primary 3.68 ms; repeats 2.09 / 2.82 / 2.70 ms. Child dirty-half loop ~54 ms wall (page fault / copy work).

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)
  • mmap vs read / O_DIRECT localhost lab

Arm B — multiprocessing fork vs spawn start cost

Empty worker that only reports VmRSS via a Queue:

Start methodstart+join p50p95child VmRSS p50
fork3.70 ms5.10 ms12500 kB
spawn61.4 ms82.6 ms16206 kB

spawn / fork ≈ 16.6× on p50 latency. Spawn also showed a fatter fresh interpreter RSS for the toy worker. On macOS/Windows, spawn (or forkserver) is the reality — Linux fork speed is not portable.


How to read these numbers

  • VmRSS ≠ unique memory after fork. Use PSS / Private_Dirty (smaps_rollup).
  • COW defers copies until write — half dirty ≈ half private growth here.
  • fork is fast to start; spawn pays import/bootstrap.
  • Container honesty: shared lab pressure (MemAvailable ~1.3 GiB) and overlay/virtio mean absolute RSS baselines include other agents — deltas are the story.

Related links:

  • fsync vs fdatasync localhost lab
  • nice / ionice CPU and disk priority lab
  • ulimit soft vs hard EMFILE lab

Pitfalls we hit (or avoided)

  1. Declaring “COW failed” because VmRSS stayed flat/high — look at Private_Dirty.
  2. Forking with threads — not tested; unsafe patterns exist.
  3. Assuming spawn is “just as fast on Linux” — 16.6× says otherwise here.
  4. Huge allocations on a memory-tight box — we used 128 MiB because MemAvailable allowed it; watch OOM.
  5. Comparing to Docker --ipc folklore without measuring PSS.

Practical checklist

  • After fork-heavy designs, monitor PSS (or cgroup memory) not only VmRSS.
  • Avoid parent writing huge shared heaps before forking workers if children will mutate them.
  • Prefer forkserver/spawn when you need safety cross-platform — budget ~tens of ms startup here.
  • Keep large read-only tables shared; mutate via messages (see pipe/tmpfile lab) when possible.
  • Re-measure under your orchestrator’s memory accounting.

Preloading big read-mostly data before fork can be a win until someone writes those pages — the Private_Dirty cliff is the bill.

Related links:

  • pipe vs tmpfile IPC localhost lab
  • zstd vs gzip vs lz4 compression localhost lab
  • SHA-256 vs BLAKE2b vs xxHash localhost lab

Versions / environment pinned

  • Python 3.13.5 (multiprocessing start methods: fork, spawn, forkserver; default fork)
  • Buffer 128 MiB bytearray, 4 KiB page touches
  • Kernel 6.12 overlay lab VM; MemAvailable ~1.3 GiB at run start


Why Shared_Dirty showed ~137 MiB

Parent had already written every page (buf[off]=1), so pages were anonymous dirty before fork. After fork they became shared dirty between parent and child (~137108 kB Shared_Dirty in the child snapshot). The child’s Private_Dirty stayed tiny (812 kB) until it wrote again.

When the child overwrote the first half, those pages faulted into private copies: Shared_Dirty fell to ~71504 kB, Private_Dirty rose to ~66416 kB. That arithmetic is the COW lecture in three numbers.

Parent VmRSS after the child exited stayed ~145 MiB — the parent still owned its copies; we did not shrink the mapping.

Methodology footnote

COW path uses bytearray so pages are anonymous and mutable (unlike a read-only bytes object). Every page is touched with a single-byte write before fork so we do not measure lazy zero-fill surprises. Child reports JSON over a pipe then _exits; parent waitpids before its own dirty loops. Start-method benches use multiprocessing.get_context(method).Process with a Queue round-trip so we time “usable worker,” not only start() returning.

Workers that only read parent data keep COW shares intact; workers that rewrite large heaps pay Private_Dirty quickly — sometimes enough that spawn + lean imports is healthier than fork + fat parent. Measure PSS under load before picking a start method for a fleet.

Verdict

Forked children shared a touched 128 MiB buffer: PSS ~70 MiB vs parent ~140 MiB with ~137 MiB Shared_Dirty. Writing half the pages cost 64 MiB Private_Dirty~~ while VmRSS barely twitched. Starting a tiny multiprocessing worker took ~3.7 ms with fork vs ~61 ms with spawn (~~16.6×). Trust COW — but trust PSS/Private_Dirty more than RSS headlines.

Evidence path on the lab box: lab-evidence/37-fork-cow-rss-vs-spawn/results/. Affiliates: 0.

fork copy-on-writecow rssmultiprocessing spawnpss private_dirtyos.forklocalhost labsrepython memory

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; 128 MiB touched bytearray; MemAvailable ~1.3 GiB on shared box. After fork child PSS ~70 MiB vs parent ~140; Shared_Dirty ~137 MiB. Child dirty half pages: Private_Dirty +65604 kB (~64 MiB), PSS +32818 kB; VmRSS barely moved (+64 kB). fork Process start+join p50 3.70 ms vs spawn 61.4 ms (~16.6x). Honest container/virtio limits. Affiliates: 0. Evidence: lab-evidence/37-fork-cow-rss-vs-spawn/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 32

    array.array vs list[int] vs bytes Lab

    A hands-on Linux localhost lab comparing memory density and numeric throughput for list[int], array.array('i'), bytearray, and memoryview.

    30 Sept 2026

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 75

    uuid.uuid4 vs uuid.uuid1: Localhost Lab

    Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What COW should look like in metrics
  3. Lab topology
  4. Arm A — COW memory after fork + half dirty
  5. Arm B — multiprocessing fork vs spawn start cost
  6. How to read these numbers
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Versions / environment pinned
  10. Why Shared\_Dirty showed \~137 MiB
  11. Methodology footnote
  12. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove