ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 99

  1. Blog

ProcessPool vs Sequential CPU: Localhost Lab

Hands-on ProcessPoolExecutor vs sequential CPU lab: real wall time and worker sweep on pure-Python sum-of-squares, measured on Linux localhost for SREs.

Aditya Challa·30 September 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — sum\_squares (p50 wall)
  5. Second workload — prng\_loop
  6. Reading it
  7. Why not ThreadPool here
  8. Spawn tax and task grain
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

When is ProcessPoolExecutor worth it over a plain sequential loop for CPU-bound Python? This lab measures wall time for 16 identical pure-Python tasks (sum-of-squares and a PRNG-style int loop) with worker sweeps 1/2/4/8 on Linux localhost.

It is not another ThreadPool vs ProcessPool GIL bake-off — that lives in the process vs thread pool GIL localhost lab. ThreadPool for I/O is covered in the ThreadPoolExecutor vs sequential localhost lab. Here the only question is: sequential → process pool, and how many workers pay off.

Related links:

  • process vs thread pool GIL localhost lab
  • threadpoolexecutor vs sequential localhost lab
  • csv reader vs split localhost lab
  • pickle vs json roundtrip localhost lab
  • hashlib md5 vs blake2b localhost lab
  • glob vs rglob vs walk localhost lab
  • json dumps compact vs indent localhost lab
  • fork COW RSS vs spawn localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5, 8 cores. concurrent.futures.ProcessPoolExecutor only. Affiliates: 0. No Docker. No ThreadPool arms in this post.

Verdict up front (sum_squares, 16 tasks × 2.5M iters): sequential ~1.461 s; 4 workers ~0.45 s (~3.24×); 8 workers ~0.363 s (~4.02×). Worker=1 is slower than sequential (~0.92×) — spawn/IPC tax with no parallelism.


Arms

ArmPattern
sequentialfor over N tasks in one process
ProcessPool w=1/2/4/8submit all tasks; wait on futures
Workloadssum_squares(2.5M), prng_loop(3M) — pure Python, holds GIL inside each worker

Lab topology

16 identical tasks · 3 rounds · p50 wall seconds
workers: 1 / 2 / 4 / 8 on 8-core box
metric: speedup = seq_p50 / pool_p50

Script: lab-evidence/80-processpool-vs-sequential/results/run_lab.py.


Lead table — sum_squares (p50 wall)

Workersp50 svs sequential
sequential1.4611.00×
11.5960.92×
20.8081.81×
40.453.24×
80.3634.02×

Second workload — prng_loop

Same shape, longer tasks (less relative spawn tax):

Workersp50 svs sequential
sequential3.5291.00×
13.6660.96×
21.8781.88×
41.0073.51×
80.7834.51×

Reading it

  • Process pools beat sequential on multi-core when each task is pure-Python CPU long enough to amortize fork/spawn.
  • w=1 is a regression — you pay pool overhead with no parallelism. Prefer sequential unless you will scale workers.
  • Scaling bends after cores are busy: ~3.2–3.5× at 4 workers, ~4.0–4.5× at 8 on this 8-core box (not linear 8× — expected).
  • Do not use ThreadPool for this class of work; GIL keeps threads near 1× (see labs 24/58).

Why not ThreadPool here

ThreadPoolExecutor on the same pure-Python loops would share one GIL and typically sit at ≤1× vs sequential. Lab 58 already measured that story for I/O sleep vs CPU threads. Lab 24 compared Process vs Thread across workload types. This post stays narrow: graduate from a for-loop to ProcessPool, pick a worker count, measure wall time.


Spawn tax and task grain

Short tasks make process pools look bad. Our tasks are multi-million-iteration loops so wall time is dominated by compute, not pickling. If your real jobs are sub-millisecond, batch them before submitting — otherwise sequential or a thread/async I/O path may win for different reasons.


Pitfalls

  • Forgetting if __name__ == "__main__" guards when the module is imported under spawn.
  • Submitting huge closures/objects — pickling cost can erase speedups.
  • Setting max_workers far above CPU count for CPU-bound work (context thrash).
  • Assuming process pools help I/O-wait (threads/async usually cheaper there).

Reproduce

python3 lab-evidence/80-processpool-vs-sequential/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box, warm, no Docker. Numbers move with core count, CPU governor, and OS scheduler. Not asyncio, not Ray/Dask. Pure-Python CPU only — C extensions that release the GIL may change the ThreadPool story (still not this post).


Takeaway

For CPU-bound pure Python, ProcessPoolExecutor with workers ≈ cores cut wall time by about ~4× vs sequential on this host. Never use w=1 as a “safe default.” Measure your task grain; keep ThreadPool for wait-bound work.

processpoolexecutorsequential vs process poolcpu-bound pythonconcurrent.futuresworker sweeplocalhost labsremultiprocessing

Lab evidence

What I found running this

Ran the Linux localhost lab on 1 Oct 2026 IST with Python 3.13.5 and 8 cores. For 16 pure-Python tasks, sequential sum_squares p50 was 1.461s; ProcessPool w=4 was 0.45s (3.24x) and w=8 was 0.363s (4.02x). w=1 regressed to 0.92x from spawn/IPC tax. Affiliates: 0.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 32

    ThreadPoolExecutor vs Sequential: I/O Lab

    Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).

    Observability & SRE · 30 Sept 2026

  • Plate 25

    ProcessPool vs ThreadPool: GIL CPU Lab with Real Numbers

    Hands-on ProcessPool vs ThreadPool lab: pure-Python threads lose to processes, hashlib threads help, and sleep I/O favors threads.

    Observability & SRE · 30 Sept 2026

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — sum\_squares (p50 wall)
  5. Second workload — prng\_loop
  6. Reading it
  7. Why not ThreadPool here
  8. Spawn tax and task grain
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove