Plate 14
as_completed vs wait: Futures Localhost Lab
Hands-on concurrent.futures as_completed vs wait lab with measured completion timing on Linux localhost.
Aditya Challa4 min read
Intro — what this post promises
You already picked ThreadPoolExecutor. How should you drain the futures — as_completed, wait(..., ALL_COMPLETED), a FIRST_COMPLETED loop, or .result() in submit order? This lab measures wall time and tasks/s on Linux localhost.
It is not ThreadPool vs sequential (lab 58) or ProcessPool vs sequential (lab 80). The pool is fixed at 8 workers; the variable is the completion API.
Related links:
- threadpoolexecutor vs sequential localhost lab
- processpoolexecutor vs sequential localhost lab
- asyncio gather vs taskgroup localhost lab
- scandir vs listdir localhost lab
- tarfile vs zipfile localhost lab
- secrets vs urandom localhost lab
- textwrap fill vs manual localhost lab
- threading event vs condition localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5, max_workers=8. Affiliates: 0. No Docker.
Verdict up front (40× sleep(5 ms) drain): all barrier-style drains clustered ~25.75–28.08 ms (~1553 tasks/s). Time-to-first: as_completed ~2.47 ms vs wait(FIRST) ~4.15 ms vs submit-order ~4.22 ms. On CPU-light tasks, as_completed ~16.4 ms beat wait(ALL) ~24.34 ms.
Arms
| Arm | Pattern |
|---|---|
as_completed | yield futures as they finish |
wait(ALL_COMPLETED) | barrier, then collect results |
wait(FIRST_COMPLETED) loop | incremental via wait |
submit-order .result() | completion order ignored |
Lab topology
Script: lab-evidence/92-futures-as-completed-vs-wait/results/run_lab.py.
Lead table — sleep drain (p50 ms)
| Workload | as_completed | wait ALL | wait FIRST loop | submit-order |
|---|---|---|---|---|
| 40×5 ms | 26.49 | 25.75 | 28.08 | 26.0 |
| 80×5 ms | 54.74 | 53.09 | 53.22 | 52.57 |
Wall ≈ sleep depth / workers — orchestration overhead is noise until you care about ordering or first result.
CPU-light drain (40 tasks)
| Arm | p50 ms | tasks/s |
|---|---|---|
| as_completed | 16.4 | 2438 |
| wait ALL | 24.34 | 1643 |
| wait FIRST loop | 23.13 | 1729 |
| submit-order | 22.11 | 1808 |
Time-to-first (40×5 ms)
| Arm | p50 ms |
|---|---|
| as_completed (first yield) | 2.47 |
| wait FIRST_COMPLETED | 4.15 |
| futs[0].result() | 4.22 |
Reading it
- Barrier completion (
wait ALL/ submit-order gather): fine when you need every result before continuing — sleep drains tied. - Incremental results: prefer
as_completed(clear API; best time-to-first here). wait(FIRST)loops work but are clunkier; similar total drain, slower first-byte in this run.- Submit-order
.result()can stall on a slow early future while later ones are done — wrong tool for streaming.
When which API
| Need | Prefer |
|---|---|
| Process results as they finish | as_completed |
| Single barrier then continue | wait(ALL_COMPLETED) |
| Select until one finishes (with timeout) | wait(FIRST_COMPLETED) |
| Preserve submit order | iterate futures + .result() |
Sleep vs CPU-light
Equal sleeps make drain APIs look identical because the wall is the sleep schedule. Uneven or CPU-bound tasks expose ordering: as_completed lets you start follow-up work sooner, which showed up as a lower total drain time on the CPU-light arm in this run.
Pitfalls
- Using submit-order
.result()when you meant completion order. - Ignoring exceptions until the end —
as_completedsurfaces them as you iterate. - Confusing this post with “threads vs processes” pool choice labs.
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box, ThreadPool only. Sleep stands in for I/O wait. Not ProcessPool, not asyncio.
Takeaway
For full drains of equal sleep tasks, as_completed ≈ wait(ALL) ≈ submit-order (~25.75 ms @40×5 ms). For first result, as_completed ~2.47 ms led. Use as_completed when incremental progress matters; use wait(ALL) for a simple barrier.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5, 8 workers. 40x sleep(5ms) drain: wait ALL 25.75ms; as_completed 26.49ms. Time-to-first: as_completed 2.47ms vs wait FIRST 4.15ms. CPU-light: as_completed 16.4ms vs wait ALL 24.34ms. Not pool-choice labs 58/80. Affiliates: 0. Evidence: lab-evidence/92-futures-as-completed-vs-wait/.
Related links
Plate 32
ThreadPoolExecutor vs Sequential: I/O Lab
Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).
Observability & SRE · 30 Sept 2026
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026