Plate 32
ThreadPoolExecutor vs Sequential: I/O Lab
Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).
Aditya Challa4 min read
Intro — what this post promises
Does concurrent.futures.ThreadPoolExecutor actually shrink wall time? This lab runs I/O-bound time.sleep tasks sequentially vs with a thread pool on Linux localhost, then contrasts a CPU-bound Python math loop — where the GIL keeps threads from helping.
Related links:
- perf_counter vs time localhost lab
- array vs list ints localhost lab
- logging vs print localhost lab
- deque vs list queue localhost lab
- lru_cache hit vs miss localhost lab
- itertools vs python loops localhost lab
- fork COW RSS vs spawn localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Sleep stands in for waiting on sockets/disk (GIL released). Affiliates: 0. Not ProcessPoolExecutor / asyncio.
Verdict up front: 40× sleep(10 ms) sequential ~411 ms; 8 workers ~55 ms (~7.4×); 40 workers ~18 ms (~23×). Same shape of CPU work with 8 threads: ~0.90× (no win — GIL).
Arms
| Arm | Pattern |
|---|---|
| sequential I/O | for _ in range(N): sleep(dt) |
| ThreadPool I/O | N submits, varying max_workers |
| sequential CPU | pure-Python sin/cos loop × 8 |
| ThreadPool CPU | same tasks on 4/8 workers |
Lab topology
Script: lab-evidence/58-threadpool-vs-sequential/results/run_lab.py.
Lead table — I/O sleep(10 ms) × 40
| Arm | p50 wall | speedup |
|---|---|---|
| sequential | 411 ms | 1.00× |
| workers=4 | 105 ms | 3.92× |
| workers=8 | 55 ms | 7.42× |
| workers=16 | 34 ms | 12.24× |
| workers=40 | 18 ms | 23.3× |
Ideal sequential ≈ 400 ms; perfect 40-way overlap ≈ 10 ms. We landed ~18 ms — overhead + scheduling, still ~23× faster than one-at-a-time.
Also: 20× 10 ms with 20 workers ~12.8×; 40× 5 ms with 40 workers ~13.5×.
CPU contrast (GIL honesty)
| Arm | p50 wall | speedup |
|---|---|---|
| sequential ×8 | 142 ms | 1.00× |
| workers=4 | 158 ms | 0.90× |
| workers=8 | 157 ms | 0.90× |
Threads did not beat sequential here (~0.9×). Pure-Python bytecode holds the GIL; use processes (or native extensions that release the GIL) for CPU fan-out.
Why sleep is a fair I/O stand-in
time.sleep releases the GIL the same way many blocking syscalls do. On this box the sequential arm tracked ~N × dt (40 × 10 ms → ~411 ms). With max_workers=N wall time collapsed toward one sleep plus pool overhead (~18 ms). That is the story for fan-out over blocking HTTP clients, DB drivers, or disk reads — not for pure-Python number crunching.
Worker sizing tip from the tables: 8 workers already delivered ~7.4× on 40 tasks; going to 40 pushed ~23×. Past “enough concurrency to cover waiters,” gains flatten and you mostly pay thread stacks.
Pitfalls
- Benchmarking CPU work with threads and concluding “threading is useless” — I/O tells a different story.
max_workers>> runnable waits — diminishing returns and memory for stacks.- Ignoring submit/join overhead on tiny sleeps — still large speedups here at 5–10 ms.
- Assuming asyncio is the only answer — ThreadPool is fine for blocking SDKs.
When to pick what
| Need | Prefer |
|---|---|
| Many blocking I/O calls | ThreadPoolExecutor |
| Pure-Python CPU fan-out | ProcessPoolExecutor / multiprocessing |
| Shared in-process state + I/O | threads |
| Structured concurrency / sockets | asyncio (different lab) |
Reproduce
Evidence: /workspace/lab-evidence/58-threadpool-vs-sequential/results/.
Closing
Threads win when work waits. On this box 40× 10 ms sleeps went from ~411 ms sequential to ~55 ms at 8 workers (~7×) and ~18 ms at 40 (~23×). The same executor on CPU math was ~0.90× — measure whether your tasks release the GIL.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. I/O: 40× sleep(10ms) sequential 411 ms; ThreadPool 8 workers 55 ms (~7.4×); 40 workers 18 ms (~23.3×). CPU math ×8: threads ~0.90× (GIL). Affiliates: 0. Evidence: lab-evidence/58-threadpool-vs-sequential/.
Related links
Plate 14
as_completed vs wait: Futures Localhost Lab
Hands-on concurrent.futures as_completed vs wait lab with measured completion timing on Linux localhost.
Observability & SRE · 30 Sept 2026
Plate 25
ProcessPool vs ThreadPool: GIL CPU Lab with Real Numbers
Hands-on ProcessPool vs ThreadPool lab: pure-Python threads lose to processes, hashlib threads help, and sleep I/O favors threads.
Observability & SRE · 30 Sept 2026
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026