Plate 99
ProcessPool vs Sequential CPU: Localhost Lab
Hands-on ProcessPoolExecutor vs sequential CPU lab: real wall time and worker sweep on pure-Python sum-of-squares, measured on Linux localhost for SREs.
Aditya Challa4 min read
Intro — what this post promises
When is ProcessPoolExecutor worth it over a plain sequential loop for CPU-bound Python? This lab measures wall time for 16 identical pure-Python tasks (sum-of-squares and a PRNG-style int loop) with worker sweeps 1/2/4/8 on Linux localhost.
It is not another ThreadPool vs ProcessPool GIL bake-off — that lives in the process vs thread pool GIL localhost lab. ThreadPool for I/O is covered in the ThreadPoolExecutor vs sequential localhost lab. Here the only question is: sequential → process pool, and how many workers pay off.
Related links:
- process vs thread pool GIL localhost lab
- threadpoolexecutor vs sequential localhost lab
- csv reader vs split localhost lab
- pickle vs json roundtrip localhost lab
- hashlib md5 vs blake2b localhost lab
- glob vs rglob vs walk localhost lab
- json dumps compact vs indent localhost lab
- fork COW RSS vs spawn localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5, 8 cores. concurrent.futures.ProcessPoolExecutor only. Affiliates: 0. No Docker. No ThreadPool arms in this post.
Verdict up front (sum_squares, 16 tasks × 2.5M iters): sequential ~1.461 s; 4 workers ~0.45 s (~3.24×); 8 workers ~0.363 s (~4.02×). Worker=1 is slower than sequential (~0.92×) — spawn/IPC tax with no parallelism.
Arms
| Arm | Pattern |
|---|---|
| sequential | for over N tasks in one process |
| ProcessPool w=1/2/4/8 | submit all tasks; wait on futures |
| Workloads | sum_squares(2.5M), prng_loop(3M) — pure Python, holds GIL inside each worker |
Lab topology
Script: lab-evidence/80-processpool-vs-sequential/results/run_lab.py.
Lead table — sum_squares (p50 wall)
| Workers | p50 s | vs sequential |
|---|---|---|
| sequential | 1.461 | 1.00× |
| 1 | 1.596 | 0.92× |
| 2 | 0.808 | 1.81× |
| 4 | 0.45 | 3.24× |
| 8 | 0.363 | 4.02× |
Second workload — prng_loop
Same shape, longer tasks (less relative spawn tax):
| Workers | p50 s | vs sequential |
|---|---|---|
| sequential | 3.529 | 1.00× |
| 1 | 3.666 | 0.96× |
| 2 | 1.878 | 1.88× |
| 4 | 1.007 | 3.51× |
| 8 | 0.783 | 4.51× |
Reading it
- Process pools beat sequential on multi-core when each task is pure-Python CPU long enough to amortize fork/spawn.
- w=1 is a regression — you pay pool overhead with no parallelism. Prefer sequential unless you will scale workers.
- Scaling bends after cores are busy: ~3.2–3.5× at 4 workers, ~4.0–4.5× at 8 on this 8-core box (not linear 8× — expected).
- Do not use ThreadPool for this class of work; GIL keeps threads near 1× (see labs 24/58).
Why not ThreadPool here
ThreadPoolExecutor on the same pure-Python loops would share one GIL and typically sit at ≤1× vs sequential. Lab 58 already measured that story for I/O sleep vs CPU threads. Lab 24 compared Process vs Thread across workload types. This post stays narrow: graduate from a for-loop to ProcessPool, pick a worker count, measure wall time.
Spawn tax and task grain
Short tasks make process pools look bad. Our tasks are multi-million-iteration loops so wall time is dominated by compute, not pickling. If your real jobs are sub-millisecond, batch them before submitting — otherwise sequential or a thread/async I/O path may win for different reasons.
Pitfalls
- Forgetting
if __name__ == "__main__"guards when the module is imported under spawn. - Submitting huge closures/objects — pickling cost can erase speedups.
- Setting
max_workersfar above CPU count for CPU-bound work (context thrash). - Assuming process pools help I/O-wait (threads/async usually cheaper there).
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box, warm, no Docker. Numbers move with core count, CPU governor, and OS scheduler. Not asyncio, not Ray/Dask. Pure-Python CPU only — C extensions that release the GIL may change the ThreadPool story (still not this post).
Takeaway
For CPU-bound pure Python, ProcessPoolExecutor with workers ≈ cores cut wall time by about ~4× vs sequential on this host. Never use w=1 as a “safe default.” Measure your task grain; keep ThreadPool for wait-bound work.
Lab evidence
What I found running this
Ran the Linux localhost lab on 1 Oct 2026 IST with Python 3.13.5 and 8 cores. For 16 pure-Python tasks, sequential sum_squares p50 was 1.461s; ProcessPool w=4 was 0.45s (3.24x) and w=8 was 0.363s (4.02x). w=1 regressed to 0.92x from spawn/IPC tax. Affiliates: 0.
Related links
Plate 32
ThreadPoolExecutor vs Sequential: I/O Lab
Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).
Observability & SRE · 30 Sept 2026
Plate 25
ProcessPool vs ThreadPool: GIL CPU Lab with Real Numbers
Hands-on ProcessPool vs ThreadPool lab: pure-Python threads lose to processes, hashlib threads help, and sleep I/O favors threads.
Observability & SRE · 30 Sept 2026
Plate 17
platform vs os.uname Inventory: Localhost Lab
Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026