ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 25

  1. Blog
  2. /Observability & SRE

ProcessPool vs ThreadPool: GIL CPU Lab with Real Numbers

Hands-on ProcessPool vs ThreadPool lab: pure-Python threads lose to processes, hashlib threads help, and sleep I/O favors threads.

Aditya Challa·30 September 2026·5 min read

Lab
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — pure-Python CPU (GIL-bound)
  5. Arm B — hashlib.sha256 (GIL released)
  6. Arm C — sleep I/O-wait
  7. When threads still win the design
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict

Intro — what this post promises

“Just add more threads” is the wrong answer for pure-Python CPU work. The GIL makes that a trap. “Always use processes” is also wrong when the work is waiting, not computing.

This is a hands-on lab with measured numbers:

  1. ThreadPoolExecutor vs ProcessPoolExecutor vs serial on 16 identical tasks.
  2. Three workloads: pure-Python math (holds the GIL), hashlib.sha256 (releases the GIL in C), and time.sleep (I/O-wait).
  3. Worker sweeps at 1 / 2 / 4 / 8 on an 8-core box.
  4. When threads help, when they hurt, and when process spawn cost shows up.

Related links:

  • Go GC + swap on a tiny VPS
  • JVM heap vs RSS vs OOM on 512 MB
  • Why your average latency graph is lying (p50 / p95 / p99)
  • Pipe vs tmpfile IPC localhost lab

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. concurrent.futures only. Local CPU + sleep — no sockets, no Docker, no GPU, no API keys. Affiliates: 0.

Verdict up front: pure-Python CPU — threads 0.82× at 4 workers (slower than serial); processes 2.41× / 3.48× . hashlib — threads 1.60× ; processes still 3.17×. Sleep I/O — threads 7.87× vs processes 6.30× .


What we compared

ArmWork per taskGIL behavior
Pure-Python CPUmath.sqrt / sin loop, 350k itersHolds the GIL
hashlib SHA-25640k × 4 KiB updates (~160 MiB hashed)Releases in C
Sleep I/O-waittime.sleep(0.05)Released while waiting

Pool API: ThreadPoolExecutor / ProcessPoolExecutor with max_workers ∈ . Metric: wall time p50 over 3 rounds for all 16 tasks.

Related links:

  • concurrent.futures — Python docs

Lab topology

16 identical tasks · serial | ThreadPool | ProcessPool
Workers: 1 / 2 / 4 / 8 (8-core box)
Arms: pure-Python CPU · hashlib.sha256 · time.sleep(0.05)
Metric: wall seconds (p50 of 3 rounds)

Arm A — pure-Python CPU (GIL-bound)

WorkersThread p50 sProcess p50 sSerial
10.5180.5560.510
20.6070.342—
40.619 (0.82×)0.211 (2.41×)—
80.6680.146 (3.48×)—

Threads never beat serial here — more workers made it worse. Processes scaled roughly with cores until spawn/overhead flattened the curve.


Arm B — hashlib.sha256 (GIL released)

WorkersThread p50 sProcess p50 sSerial
11.7991.7551.819
21.2500.941—
41.140 (1.60×)0.574 (3.17×)—
81.580 (regressed)0.509 (3.57×)—

C extensions that release the GIL let threads help — but processes still won big, and thread w=8 regressed vs w=4 on this shared box (contention / scheduling noise).


Arm C — sleep I/O-wait

WorkersThread p50 sProcess p50 sSerial
10.8050.8150.802
40.2050.215—
80.102 (7.87×)0.127 (6.30×)—

When tasks mostly wait, threads are enough — and slightly cheaper than forking eight workers for 50 ms sleeps.

Related links:

  • Pipe vs tmpfile IPC localhost lab
  • epoll vs select FD_SETSIZE lab

When threads still win the design

  • I/O-bound / wait-bound work (network, disk, sleep, locks you release around).
  • C extensions that release the GIL (hashlib, many crypto/compress codecs) — measure; do not assume.
  • Shared in-memory state you do not want to pickle across processes.
  • Tiny tasks where process spawn + IPC dwarfs the work (Arm C’s process tax).

Reach for processes when the hotspot is pure Python (or any code that holds the GIL for long stretches).


Pitfalls we hit (or avoided)

  1. “Threads = parallelism” for pure Python — Arm A is the counterexample (0.82×).
  2. Assuming hashlib needs processes — threads helped (1.6×); processes still better if CPU-bound hashing dominates.
  3. More workers always better — thread hashlib w=8 regressing vs w=4.
  4. Ignoring process start cost — Arm C: threads edged processes at 8 workers.
  5. Calling this an asyncio result — this lab is thread/process pools, not the event loop.

Practical checklist

  • Profile: is the hot path pure Python or C / I/O?
  • Pure-Python CPU → ProcessPoolExecutor (or another process model); do not expect thread speedups.
  • Wait-heavy work → threads (or asyncio) first; measure before paying process overhead.
  • Report workload type + worker count + wall p50 with every speedup claim.
  • Cap workers near core count for CPU arms; I/O arms may want more — still measure.

Verdict

On this 8-core box with 16 tasks: pure-Python CPU — threads lost (0.82× workers); processes 2.41× / 3.48× . hashlib — threads 1.60× , processes 3.17×. Sleep I/O — threads 7.87× edged processes 6.30× at 8 workers. Match the pool to the GIL story, not the folklore.

Evidence path on the lab box: lab-evidence/24-process-vs-thread-pool/results/. Affiliates: 0.

processpoolexecutor vs threadpoolexecutorpython gil parallelismthread pool cpu boundmultiprocessing vs threadinghashlib gil releaselocalhost labsreconcurrent.futures

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5 concurrent.futures. 16 tasks on 8-core box. Pure-Python CPU: serial 0.510s; thread w4 0.619s (0.82× slower); process w4 0.211s (2.41×); process w8 0.146s (3.48×). hashlib.sha256: serial 1.819s; thread w4 1.140s (1.60×); process w4 0.574s (3.17×); thread w8 regresses to 1.580s. sleep(0.05): serial 0.802s; thread w8 0.102s (7.87×); process w8 0.127s (6.30×). No Docker. Affiliates: 0. Evidence: lab-evidence/24-process-vs-thread-pool/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 32

    ThreadPoolExecutor vs Sequential: I/O Lab

    Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).

    Observability & SRE · 30 Sept 2026

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — pure-Python CPU (GIL-bound)
  5. Arm B — hashlib.sha256 (GIL released)
  6. Arm C — sleep I/O-wait
  7. When threads still win the design
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove