ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 32

  1. Blog
  2. /Observability & SRE

ThreadPoolExecutor vs Sequential: I/O Lab

Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — I/O sleep(10 ms) × 40
  5. CPU contrast (GIL honesty)
  6. Why sleep is a fair I/O stand-in
  7. Pitfalls
  8. When to pick what
  9. Reproduce
  10. Closing

Intro — what this post promises

Does concurrent.futures.ThreadPoolExecutor actually shrink wall time? This lab runs I/O-bound time.sleep tasks sequentially vs with a thread pool on Linux localhost, then contrasts a CPU-bound Python math loop — where the GIL keeps threads from helping.

Related links:

  • perf_counter vs time localhost lab
  • array vs list ints localhost lab
  • logging vs print localhost lab
  • deque vs list queue localhost lab
  • lru_cache hit vs miss localhost lab
  • itertools vs python loops localhost lab
  • fork COW RSS vs spawn localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Sleep stands in for waiting on sockets/disk (GIL released). Affiliates: 0. Not ProcessPoolExecutor / asyncio.

Verdict up front: 40× sleep(10 ms) sequential ~411 ms; 8 workers ~55 ms (~7.4×); 40 workers ~18 ms (~23×). Same shape of CPU work with 8 threads: ~0.90× (no win — GIL).


Arms

ArmPattern
sequential I/Ofor _ in range(N): sleep(dt)
ThreadPool I/ON submits, varying max_workers
sequential CPUpure-Python sin/cos loop × 8
ThreadPool CPUsame tasks on 4/8 workers

Lab topology

I/O: N in {20,40}, dt in {5ms,10ms}
CPU: 8 tasks of ~200k float ops each
Metric: p50 wall; speedup = seq_p50 / pool_p50

Script: lab-evidence/58-threadpool-vs-sequential/results/run_lab.py.


Lead table — I/O sleep(10 ms) × 40

Armp50 wallspeedup
sequential411 ms1.00×
workers=4105 ms3.92×
workers=855 ms7.42×
workers=1634 ms12.24×
workers=4018 ms23.3×

Ideal sequential ≈ 400 ms; perfect 40-way overlap ≈ 10 ms. We landed ~18 ms — overhead + scheduling, still ~23× faster than one-at-a-time.

Also: 20× 10 ms with 20 workers ~12.8×; 40× 5 ms with 40 workers ~13.5×.


CPU contrast (GIL honesty)

Armp50 wallspeedup
sequential ×8142 ms1.00×
workers=4158 ms0.90×
workers=8157 ms0.90×

Threads did not beat sequential here (~0.9×). Pure-Python bytecode holds the GIL; use processes (or native extensions that release the GIL) for CPU fan-out.


Why sleep is a fair I/O stand-in

time.sleep releases the GIL the same way many blocking syscalls do. On this box the sequential arm tracked ~N × dt (40 × 10 ms → ~411 ms). With max_workers=N wall time collapsed toward one sleep plus pool overhead (~18 ms). That is the story for fan-out over blocking HTTP clients, DB drivers, or disk reads — not for pure-Python number crunching.

Worker sizing tip from the tables: 8 workers already delivered ~7.4× on 40 tasks; going to 40 pushed ~23×. Past “enough concurrency to cover waiters,” gains flatten and you mostly pay thread stacks.


Pitfalls

  1. Benchmarking CPU work with threads and concluding “threading is useless” — I/O tells a different story.
  2. max_workers >> runnable waits — diminishing returns and memory for stacks.
  3. Ignoring submit/join overhead on tiny sleeps — still large speedups here at 5–10 ms.
  4. Assuming asyncio is the only answer — ThreadPool is fine for blocking SDKs.

When to pick what

NeedPrefer
Many blocking I/O callsThreadPoolExecutor
Pure-Python CPU fan-outProcessPoolExecutor / multiprocessing
Shared in-process state + I/Othreads
Structured concurrency / socketsasyncio (different lab)

Reproduce

python3 lab-evidence/58-threadpool-vs-sequential/results/run_lab.py

Evidence: /workspace/lab-evidence/58-threadpool-vs-sequential/results/.


Closing

Threads win when work waits. On this box 40× 10 ms sleeps went from ~411 ms sequential to ~55 ms at 8 workers (~7×) and ~18 ms at 40 (~23×). The same executor on CPU math was ~0.90× — measure whether your tasks release the GIL.

threadpoolexecutorconcurrent.futuresthreading i/ogilpython threadslocalhost labsreparallelism

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. I/O: 40× sleep(10ms) sequential 411 ms; ThreadPool 8 workers 55 ms (~7.4×); 40 workers 18 ms (~23.3×). CPU math ×8: threads ~0.90× (GIL). Affiliates: 0. Evidence: lab-evidence/58-threadpool-vs-sequential/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 14

    as_completed vs wait: Futures Localhost Lab

    Hands-on concurrent.futures as_completed vs wait lab with measured completion timing on Linux localhost.

    Observability & SRE · 30 Sept 2026

  • Plate 25

    ProcessPool vs ThreadPool: GIL CPU Lab with Real Numbers

    Hands-on ProcessPool vs ThreadPool lab: pure-Python threads lose to processes, hashlib threads help, and sleep I/O favors threads.

    Observability & SRE · 30 Sept 2026

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — I/O sleep(10 ms) × 40
  5. CPU contrast (GIL honesty)
  6. Why sleep is a fair I/O stand-in
  7. Pitfalls
  8. When to pick what
  9. Reproduce
  10. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove