Plate 52
asyncio.gather vs TaskGroup: Localhost Lab
Aditya Challa4 min read
Intro — what this post promises
How should you orchestrate many concurrent awaits — asyncio.gather, asyncio.TaskGroup, or a plain sequential for/await loop? This lab measures wall time for N I/O-ish asyncio.sleep tasks on Linux localhost.
It is not another asyncio-vs-threads bake-off (see asyncio vs threads localhost lab) and not a kernel context-switch study. The question is narrowly: gather vs TaskGroup vs sequential awaits for overlapping waits.
Related links:
- asyncio vs threads localhost lab
- processpoolexecutor vs sequential localhost lab
- threadpoolexecutor vs sequential localhost lab
- mmap vs read scan localhost lab
- csv reader vs split localhost lab
- pickle vs json roundtrip localhost lab
- hashlib md5 vs blake2b localhost lab
- glob vs rglob vs walk localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. asyncio.sleep stands in for overlapping I/O waits (GIL released while sleeping). Affiliates: 0. No Docker. No ThreadPool arms.
Verdict up front (40× sleep(10 ms)): sequential ~0.415 s; gather ~0.0109 s (~38.15×); TaskGroup ~0.011 s (~37.87×). At 100× sleep(5 ms) both concurrent arms stay near the ~5 ms ideal (~79.3–83.63× vs sequential). gather ≈ TaskGroup for happy-path throughput.
Arms
| Arm | Pattern |
|---|---|
| sequential | for _ in range(N): await sleep(dt) |
asyncio.gather | await gather(*(sleep(dt) for _ in range(N))) |
asyncio.TaskGroup | async with TaskGroup() as tg: tg.create_task(...) |
gather + return_exceptions=True | same overlap; exception policy differs |
Lab topology
Script: lab-evidence/82-asyncio-gather-vs-taskgroup/results/run_lab.py.
Lead table — 40× 10 ms (p50)
| Arm | p50 s | vs sequential |
|---|---|---|
| sequential | 0.415 | 1.00× |
| gather | 0.0109 | 38.15× |
| TaskGroup | 0.011 | 37.87× |
| gather + return_exceptions | 0.0109 | 37.97× |
Ideal concurrent wall ≈ 0.010 s; both gather and TaskGroup land within ~10% of that.
Scale check — 100× 5 ms
| Arm | p50 s | vs sequential |
|---|---|---|
| sequential | 0.53 | 1.00× |
| gather | 0.0067 | 79.3× |
| TaskGroup | 0.0063 | 83.63× |
Reading it
- Sequential awaits serialize waits — wall ≈ N×dt. Never do this for independent I/O.
- gather and TaskGroup are tied for this happy-path sleep fan-out (~38× at 40×10 ms; ~80× at 100×5 ms).
return_exceptions=Truedid not change throughput here (no failures); it changes error policy, not speed.- Prefer TaskGroup when you want structured concurrency (exceptions cancel siblings); prefer gather when you need a result list API / older patterns.
gather vs TaskGroup (behavior, not speed)
On success they look the same in wall time. They diverge when tasks fail: TaskGroup raises an ExceptionGroup and cancels siblings; gather can collect exceptions with return_exceptions=True or fail-fast on the first error. Pick for semantics, not microseconds.
Why sleep is a fair proxy here
asyncio.sleep yields the event loop like waiting on sockets/timers. This post is about scheduling many overlapping awaits, not syscall fidelity. Real TCP would add connection setup noise (lab 28 already covered asyncio vs threads on localhost echo).
Ideal wall vs measured
Ideal concurrent wall is one dt. At 40×10 ms, gather/TaskGroup sat ~10.8–11.0 ms — about 1.1× ideal — the gap is task creation + loop scheduling. Sequential paid the full N×dt (~415 ms). If your “async” code still looks like sequential awaits, you will never see that ~38× cliff.
Pitfalls
- Awaiting tasks one-by-one in a
forloop — the classic “async but serial” bug. - Creating tasks without awaiting gather/TaskGroup — orphaned tasks / “Task was destroyed but it is pending”.
- Assuming TaskGroup is “slower” without measuring — here it matched gather.
- Mixing CPU-bound work into the loop without an executor (GIL / stalls).
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box. Sleep-only I/O proxy. No uvloop. Python 3.11+ TaskGroup. Numbers move with timer resolution and load.
Takeaway
For overlapping waits, asyncio.gather and TaskGroup both crush sequential awaits (~38.15× at 40×10 ms; ~83.63× at 100×5 ms) and stay near the ideal one-delay wall. Choose TaskGroup for structured failure/cancel; gather for list results / return_exceptions. Stop writing sequential await loops for independent I/O.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Ran Python 3.13.5 on Linux localhost using asyncio.sleep as an overlapping I/O proxy. At 40 x sleep(10 ms), sequential was 0.415 s, gather 0.0109 s (38.15x), and TaskGroup 0.011 s (37.87x). At 100 x sleep(5 ms), gather was 79.3x and TaskGroup 83.63x. Gather and TaskGroup were effectively tied on happy-path throughput; no ThreadPool arms, Docker, or affiliate links were used.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026