ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 14

  1. Blog
  2. /Observability & SRE

as_completed vs wait: Futures Localhost Lab

Hands-on concurrent.futures as_completed vs wait lab with measured completion timing on Linux localhost.

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — sleep drain (p50 ms)
  5. CPU-light drain (40 tasks)
  6. Time-to-first (40×5 ms)
  7. Reading it
  8. When which API
  9. Sleep vs CPU-light
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway

Intro — what this post promises

You already picked ThreadPoolExecutor. How should you drain the futures — as_completed, wait(..., ALL_COMPLETED), a FIRST_COMPLETED loop, or .result() in submit order? This lab measures wall time and tasks/s on Linux localhost.

It is not ThreadPool vs sequential (lab 58) or ProcessPool vs sequential (lab 80). The pool is fixed at 8 workers; the variable is the completion API.

Related links:

  • threadpoolexecutor vs sequential localhost lab
  • processpoolexecutor vs sequential localhost lab
  • asyncio gather vs taskgroup localhost lab
  • scandir vs listdir localhost lab
  • tarfile vs zipfile localhost lab
  • secrets vs urandom localhost lab
  • textwrap fill vs manual localhost lab
  • threading event vs condition localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5, max_workers=8. Affiliates: 0. No Docker.

Verdict up front (40× sleep(5 ms) drain): all barrier-style drains clustered ~25.75–28.08 ms (~1553 tasks/s). Time-to-first: as_completed ~2.47 ms vs wait(FIRST) ~4.15 ms vs submit-order ~4.22 ms. On CPU-light tasks, as_completed ~16.4 ms beat wait(ALL) ~24.34 ms.


Arms

ArmPattern
as_completedyield futures as they finish
wait(ALL_COMPLETED)barrier, then collect results
wait(FIRST_COMPLETED) loopincremental via wait
submit-order .result()completion order ignored

Lab topology

workers=8 · rounds=7 · p50 wall for drain phase
sleep: 40× / 80× 5 ms
cpu_light: 40× sum-of-squares(50k)
also: time-to-first-result on 40×5 ms

Script: lab-evidence/92-futures-as-completed-vs-wait/results/run_lab.py.


Lead table — sleep drain (p50 ms)

Workloadas_completedwait ALLwait FIRST loopsubmit-order
40×5 ms26.4925.7528.0826.0
80×5 ms54.7453.0953.2252.57

Wall ≈ sleep depth / workers — orchestration overhead is noise until you care about ordering or first result.


CPU-light drain (40 tasks)

Armp50 mstasks/s
as_completed16.42438
wait ALL24.341643
wait FIRST loop23.131729
submit-order22.111808

Time-to-first (40×5 ms)

Armp50 ms
as_completed (first yield)2.47
wait FIRST_COMPLETED4.15
futs[0].result()4.22

Reading it

  • Barrier completion (wait ALL / submit-order gather): fine when you need every result before continuing — sleep drains tied.
  • Incremental results: prefer as_completed (clear API; best time-to-first here).
  • wait(FIRST) loops work but are clunkier; similar total drain, slower first-byte in this run.
  • Submit-order .result() can stall on a slow early future while later ones are done — wrong tool for streaming.

When which API

NeedPrefer
Process results as they finishas_completed
Single barrier then continuewait(ALL_COMPLETED)
Select until one finishes (with timeout)wait(FIRST_COMPLETED)
Preserve submit orderiterate futures + .result()

Sleep vs CPU-light

Equal sleeps make drain APIs look identical because the wall is the sleep schedule. Uneven or CPU-bound tasks expose ordering: as_completed lets you start follow-up work sooner, which showed up as a lower total drain time on the CPU-light arm in this run.


Pitfalls

  • Using submit-order .result() when you meant completion order.
  • Ignoring exceptions until the end — as_completed surfaces them as you iterate.
  • Confusing this post with “threads vs processes” pool choice labs.

Reproduce

python3 lab-evidence/92-futures-as-completed-vs-wait/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box, ThreadPool only. Sleep stands in for I/O wait. Not ProcessPool, not asyncio.


Takeaway

For full drains of equal sleep tasks, as_completed ≈ wait(ALL) ≈ submit-order (~25.75 ms @40×5 ms). For first result, as_completed ~2.47 ms led. Use as_completed when incremental progress matters; use wait(ALL) for a simple barrier.

as_completedconcurrent.futures waitall_completedfirst_completedthreadpoolexecutorlocalhost labsrepython futures

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5, 8 workers. 40x sleep(5ms) drain: wait ALL 25.75ms; as_completed 26.49ms. Time-to-first: as_completed 2.47ms vs wait FIRST 4.15ms. CPU-light: as_completed 16.4ms vs wait ALL 24.34ms. Not pool-choice labs 58/80. Affiliates: 0. Evidence: lab-evidence/92-futures-as-completed-vs-wait/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 32

    ThreadPoolExecutor vs Sequential: I/O Lab

    Hands-on ThreadPoolExecutor vs sequential lab: real wall-time speedup for I/O sleep workloads plus GIL CPU contrast, measured on Linux localhost (lab).

    Observability & SRE · 30 Sept 2026

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — sleep drain (p50 ms)
  5. CPU-light drain (40 tasks)
  6. Time-to-first (40×5 ms)
  7. Reading it
  8. When which API
  9. Sleep vs CPU-light
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove