Plate 25
ProcessPool vs ThreadPool: GIL CPU Lab with Real Numbers
Hands-on ProcessPool vs ThreadPool lab: pure-Python threads lose to processes, hashlib threads help, and sleep I/O favors threads.
Aditya Challa5 min read
Intro — what this post promises
“Just add more threads” is the wrong answer for pure-Python CPU work. The GIL makes that a trap. “Always use processes” is also wrong when the work is waiting, not computing.
This is a hands-on lab with measured numbers:
ThreadPoolExecutorvsProcessPoolExecutorvs serial on 16 identical tasks.- Three workloads: pure-Python math (holds the GIL),
hashlib.sha256(releases the GIL in C), andtime.sleep(I/O-wait). - Worker sweeps at 1 / 2 / 4 / 8 on an 8-core box.
- When threads help, when they hurt, and when process spawn cost shows up.
Related links:
- Go GC + swap on a tiny VPS
- JVM heap vs RSS vs OOM on 512 MB
- Why your average latency graph is lying (p50 / p95 / p99)
- Pipe vs tmpfile IPC localhost lab
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. concurrent.futures only. Local CPU + sleep — no sockets, no Docker, no GPU, no API keys. Affiliates: 0.
Verdict up front: pure-Python CPU — threads 0.82× at 4 workers (slower than serial); processes 2.41× / 3.48× . hashlib — threads 1.60× ; processes still 3.17×. Sleep I/O — threads 7.87× vs processes 6.30× .
What we compared
| Arm | Work per task | GIL behavior |
|---|---|---|
| Pure-Python CPU | math.sqrt / sin loop, 350k iters | Holds the GIL |
| hashlib SHA-256 | 40k × 4 KiB updates (~160 MiB hashed) | Releases in C |
| Sleep I/O-wait | time.sleep(0.05) | Released while waiting |
Pool API: ThreadPoolExecutor / ProcessPoolExecutor with max_workers ∈ . Metric: wall time p50 over 3 rounds for all 16 tasks.
Related links:
Lab topology
Arm A — pure-Python CPU (GIL-bound)
| Workers | Thread p50 s | Process p50 s | Serial |
|---|---|---|---|
| 1 | 0.518 | 0.556 | 0.510 |
| 2 | 0.607 | 0.342 | — |
| 4 | 0.619 (0.82×) | 0.211 (2.41×) | — |
| 8 | 0.668 | 0.146 (3.48×) | — |
Threads never beat serial here — more workers made it worse. Processes scaled roughly with cores until spawn/overhead flattened the curve.
Arm B — hashlib.sha256 (GIL released)
| Workers | Thread p50 s | Process p50 s | Serial |
|---|---|---|---|
| 1 | 1.799 | 1.755 | 1.819 |
| 2 | 1.250 | 0.941 | — |
| 4 | 1.140 (1.60×) | 0.574 (3.17×) | — |
| 8 | 1.580 (regressed) | 0.509 (3.57×) | — |
C extensions that release the GIL let threads help — but processes still won big, and thread w=8 regressed vs w=4 on this shared box (contention / scheduling noise).
Arm C — sleep I/O-wait
| Workers | Thread p50 s | Process p50 s | Serial |
|---|---|---|---|
| 1 | 0.805 | 0.815 | 0.802 |
| 4 | 0.205 | 0.215 | — |
| 8 | 0.102 (7.87×) | 0.127 (6.30×) | — |
When tasks mostly wait, threads are enough — and slightly cheaper than forking eight workers for 50 ms sleeps.
Related links:
When threads still win the design
- I/O-bound / wait-bound work (network, disk, sleep, locks you release around).
- C extensions that release the GIL (
hashlib, many crypto/compress codecs) — measure; do not assume. - Shared in-memory state you do not want to pickle across processes.
- Tiny tasks where process spawn + IPC dwarfs the work (Arm C’s process tax).
Reach for processes when the hotspot is pure Python (or any code that holds the GIL for long stretches).
Pitfalls we hit (or avoided)
- “Threads = parallelism” for pure Python — Arm A is the counterexample (0.82×).
- Assuming hashlib needs processes — threads helped (1.6×); processes still better if CPU-bound hashing dominates.
- More workers always better — thread hashlib w=8 regressing vs w=4.
- Ignoring process start cost — Arm C: threads edged processes at 8 workers.
- Calling this an asyncio result — this lab is thread/process pools, not the event loop.
Practical checklist
- Profile: is the hot path pure Python or C / I/O?
- Pure-Python CPU →
ProcessPoolExecutor(or another process model); do not expect thread speedups. - Wait-heavy work → threads (or asyncio) first; measure before paying process overhead.
- Report workload type + worker count + wall p50 with every speedup claim.
- Cap workers near core count for CPU arms; I/O arms may want more — still measure.
Verdict
On this 8-core box with 16 tasks: pure-Python CPU — threads lost (0.82× workers); processes 2.41× / 3.48× . hashlib — threads 1.60× , processes 3.17×. Sleep I/O — threads 7.87× edged processes 6.30× at 8 workers. Match the pool to the GIL story, not the folklore.
Evidence path on the lab box: lab-evidence/24-process-vs-thread-pool/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5 concurrent.futures. 16 tasks on 8-core box. Pure-Python CPU: serial 0.510s; thread w4 0.619s (0.82× slower); process w4 0.211s (2.41×); process w8 0.146s (3.48×). hashlib.sha256: serial 1.819s; thread w4 1.140s (1.60×); process w4 0.574s (3.17×); thread w8 regresses to 1.580s. sleep(0.05): serial 0.802s; thread w8 0.102s (7.87×); process w8 0.127s (6.30×). No Docker. Affiliates: 0. Evidence: lab-evidence/24-process-vs-thread-pool/.
Related links
Plate 32
ThreadPoolExecutor vs Sequential: I/O Lab
Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).
Observability & SRE · 30 Sept 2026
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026