ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 52

  1. Blog
  2. /Observability & SRE

asyncio.gather vs TaskGroup: Localhost Lab

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 40× 10 ms (p50)
  5. Scale check — 100× 5 ms
  6. Reading it
  7. gather vs TaskGroup (behavior, not speed)
  8. Why sleep is a fair proxy here
  9. Ideal wall vs measured
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway

Intro — what this post promises

How should you orchestrate many concurrent awaits — asyncio.gather, asyncio.TaskGroup, or a plain sequential for/await loop? This lab measures wall time for N I/O-ish asyncio.sleep tasks on Linux localhost.

It is not another asyncio-vs-threads bake-off (see asyncio vs threads localhost lab) and not a kernel context-switch study. The question is narrowly: gather vs TaskGroup vs sequential awaits for overlapping waits.

Related links:

  • asyncio vs threads localhost lab
  • processpoolexecutor vs sequential localhost lab
  • threadpoolexecutor vs sequential localhost lab
  • mmap vs read scan localhost lab
  • csv reader vs split localhost lab
  • pickle vs json roundtrip localhost lab
  • hashlib md5 vs blake2b localhost lab
  • glob vs rglob vs walk localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. asyncio.sleep stands in for overlapping I/O waits (GIL released while sleeping). Affiliates: 0. No Docker. No ThreadPool arms.

Verdict up front (40× sleep(10 ms)): sequential ~0.415 s; gather ~0.0109 s (~38.15×); TaskGroup ~0.011 s (~37.87×). At 100× sleep(5 ms) both concurrent arms stay near the ~5 ms ideal (~79.3–83.63× vs sequential). gather ≈ TaskGroup for happy-path throughput.


Arms

ArmPattern
sequentialfor _ in range(N): await sleep(dt)
asyncio.gatherawait gather(*(sleep(dt) for _ in range(N)))
asyncio.TaskGroupasync with TaskGroup() as tg: tg.create_task(...)
gather + return_exceptions=Truesame overlap; exception policy differs

Lab topology

N in {20,40,100} · dt in {5ms,10ms} · rounds=5 · p50 wall
metric: speedup = seq_p50 / arm_p50
ideal concurrent wall ≈ dt (not N×dt)

Script: lab-evidence/82-asyncio-gather-vs-taskgroup/results/run_lab.py.


Lead table — 40× 10 ms (p50)

Armp50 svs sequential
sequential0.4151.00×
gather0.010938.15×
TaskGroup0.01137.87×
gather + return_exceptions0.010937.97×

Ideal concurrent wall ≈ 0.010 s; both gather and TaskGroup land within ~10% of that.


Scale check — 100× 5 ms

Armp50 svs sequential
sequential0.531.00×
gather0.006779.3×
TaskGroup0.006383.63×

Reading it

  • Sequential awaits serialize waits — wall ≈ N×dt. Never do this for independent I/O.
  • gather and TaskGroup are tied for this happy-path sleep fan-out (~38× at 40×10 ms; ~80× at 100×5 ms).
  • return_exceptions=True did not change throughput here (no failures); it changes error policy, not speed.
  • Prefer TaskGroup when you want structured concurrency (exceptions cancel siblings); prefer gather when you need a result list API / older patterns.

gather vs TaskGroup (behavior, not speed)

On success they look the same in wall time. They diverge when tasks fail: TaskGroup raises an ExceptionGroup and cancels siblings; gather can collect exceptions with return_exceptions=True or fail-fast on the first error. Pick for semantics, not microseconds.


Why sleep is a fair proxy here

asyncio.sleep yields the event loop like waiting on sockets/timers. This post is about scheduling many overlapping awaits, not syscall fidelity. Real TCP would add connection setup noise (lab 28 already covered asyncio vs threads on localhost echo).


Ideal wall vs measured

Ideal concurrent wall is one dt. At 40×10 ms, gather/TaskGroup sat ~10.8–11.0 ms — about 1.1× ideal — the gap is task creation + loop scheduling. Sequential paid the full N×dt (~415 ms). If your “async” code still looks like sequential awaits, you will never see that ~38× cliff.


Pitfalls

  • Awaiting tasks one-by-one in a for loop — the classic “async but serial” bug.
  • Creating tasks without awaiting gather/TaskGroup — orphaned tasks / “Task was destroyed but it is pending”.
  • Assuming TaskGroup is “slower” without measuring — here it matched gather.
  • Mixing CPU-bound work into the loop without an executor (GIL / stalls).

Reproduce

python3 lab-evidence/82-asyncio-gather-vs-taskgroup/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box. Sleep-only I/O proxy. No uvloop. Python 3.11+ TaskGroup. Numbers move with timer resolution and load.


Takeaway

For overlapping waits, asyncio.gather and TaskGroup both crush sequential awaits (~38.15× at 40×10 ms; ~83.63× at 100×5 ms) and stay near the ideal one-delay wall. Choose TaskGroup for structured failure/cancel; gather for list results / return_exceptions. Stop writing sequential await loops for independent I/O.

asyncio.gatherasyncio.taskgroupsequential awaitsstructured concurrencypython asynciolocalhost labsreorchestration

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Ran Python 3.13.5 on Linux localhost using asyncio.sleep as an overlapping I/O proxy. At 40 x sleep(10 ms), sequential was 0.415 s, gather 0.0109 s (38.15x), and TaskGroup 0.011 s (37.87x). At 100 x sleep(5 ms), gather was 79.3x and TaskGroup 83.63x. Gather and TaskGroup were effectively tied on happy-path throughput; no ThreadPool arms, Docker, or affiliate links were used.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 40× 10 ms (p50)
  5. Scale check — 100× 5 ms
  6. Reading it
  7. gather vs TaskGroup (behavior, not speed)
  8. Why sleep is a fair proxy here
  9. Ideal wall vs measured
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove