Plate 80
Asyncio vs Threads: I/O Concurrency Lab with Real Numbers
Hands-on asyncio vs ThreadPool lab: sleep 58.7× vs 44×; delayed TCP 13.8× vs serial; fast-echo trap loses; CPU flat. Real localhost numbers, no Docker.
Aditya Challa6 min read
On this page
- Intro — what this post promises
- What we compared
- Lab topology
- Arm A — concurrent sleep (the textbook win)
- Arm B — delayed TCP echo (5 ms think time)
- Arm C — instant echo (the overhead trap)
- Arm D — mixed sleep + slow TCP
- Arm E — pure-Python CPU (GIL control)
- When to pick which
- Pitfalls we hit (or avoided)
- Practical checklist
- Verdict
Intro — what this post promises
Asyncio and threads both overlap waiting work. They are not interchangeable: one multiplexes on a single event loop; the other pays thread stacks and the GIL tax when you accidentally do CPU.
This is a hands-on lab with measured numbers:
- Concurrent
sleep— ideal wall ≈ one delay, not N×delay. - Delayed localhost TCP echo — where concurrency actually wins.
- Instant echo — the overhead trap where serial wins.
- Mixed sleep + slow TCP on one loop vs a thread pool.
- Pure-Python CPU control — threads do not parallelize it.
Related links:
- ProcessPool vs ThreadPool GIL localhost lab
- Pipe vs tmpfile IPC localhost lab
- epoll vs select FD_SETSIZE localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 asyncio + ThreadPoolExecutor against a 127.0.0.1 echo (instant and 5 ms delayed). No Docker. No GPU. No API keys. Affiliates: 0.
Verdict up front: sleep asyncio ~59× / threads ~44× vs ideal serial; delayed TCP asyncio ~14× vs serial (~2.4× vs threads w32); instant echo concurrent arms lost; pure-Python CPU threads ~0.92×.
What we compared
| Arm | Model | Metric |
|---|---|---|
| A | serial / threads / asyncio sleep | wall vs ideal N×delay |
| B | delayed TCP echo (5 ms server think) | wall RPS + speedup |
| C | instant TCP echo (overhead trap) | RPS vs serial |
| D | mixed sleep + delayed TCP | wall speedup |
| E | pure-Python CPU | threads vs serial (GIL) |
Related links:
Lab topology
Arm A — concurrent sleep (the textbook win)
n=64, delay=20 ms, ideal serial = 1.280 s.
| Model | Wall s | Speedup vs ideal |
|---|---|---|
| serial | 1.301 | ~1.0× |
| ThreadPool | 0.029 | 44.4× |
| asyncio | 0.022 | 58.7× |
Both models collapse wall time toward one delay. Asyncio edged threads here: no per-task OS thread.
Arm B — delayed TCP echo (5 ms think time)
n=64 connect-per 64 B echo. Ideal serial wait ≈ 0.32 s (delay alone).
| Model | Wall s | RPS | vs serial |
|---|---|---|---|
| serial | 0.393 | 163 | 1.0× |
| threads w=8 | 0.058 | 1107 | 6.8× |
| threads w=16 | 0.042 | 1515 | 9.3× |
| threads w=32 | 0.067 | 954 | 5.9× |
| threads w=64 | 0.051 | 1247 | 7.7× |
| asyncio | 0.029 | 2245 | 13.8× |
When each request waits on the peer, overlapping waits is the whole game. Asyncio beat the best thread pool by ~2.4× vs w32 on this box (thread-count sweet spot was w16, not “max workers = n”).
Related links:
Arm C — instant echo (the overhead trap)
Same shape, 0 ms server delay, n=200:
| Model | RPS | vs serial |
|---|---|---|
| serial | 2658 | 1.0× |
| threads w32 | 2016 | 0.76× |
| asyncio | 1337 | 0.50× |
Localhost connect+echo is so cheap that thread/task setup loses to a tight serial loop. Concurrency is not free — measure the wait you are overlapping.
Arm D — mixed sleep + slow TCP
32× sleep(20 ms) + 32× delayed TCP:
| Model | Wall s | vs serial |
|---|---|---|
| serial | 0.835 | 1.0× |
| ThreadPool | 0.029 | 28.6× |
| asyncio | 0.021 | 39.1× |
One event loop (or one pool) overlaps both wait types. This is the “many sockets + timers” shape asyncio is for.
Arm E — pure-Python CPU (GIL control)
8 tasks of tight pure-Python arithmetic:
| Model | Wall s | vs serial |
|---|---|---|
| serial | 0.047 | 1.0× |
| ThreadPool(8) | 0.051 | 0.92× |
Threads did not speed this up. For CPU-bound pure Python, use processes (see the GIL pool lab) — not more threads, and not await on hot loops.
Related links:
When to pick which
| Situation | Prefer |
|---|---|
| Many waits (sockets, sleeps, timers) | asyncio (or threads if the ecosystem is sync-only) |
| Blocking sync library you cannot rewrite | ThreadPoolExecutor off the loop |
| Pure-Python / GIL-bound CPU | ProcessPool (not this post’s winners) |
| Microsecond localhost RPC with no wait | Measure — serial may win (Arm C) |
Pitfalls we hit (or avoided)
- Benchmarking only instant localhost echo — Arm C lies about I/O concurrency value.
- Assuming more thread workers always help — w16 beat w32/w64 on Arm B.
- Putting CPU work on the event loop — serializes everything behind it.
- Treating this as uvloop / Trio / gevent — we measured stdlib asyncio + threads only.
- Confusing with ProcessPool CPU scaling — different problem; different lab.
Practical checklist
- Name the wait you are overlapping (delay ms, RTT, disk). If wait ≈ 0, concurrency may lose.
- Report wall + RPS + worker/concurrency + serial baseline together.
- Cap thread workers; sweep (8/16/32) instead of
workers=n_tasksby habit. - Keep CPU off the loop; hand it to processes or native extensions that release the GIL.
- Re-run under your real client library — wrappers add their own tax.
Verdict
On this box: concurrent sleep favored asyncio (~59×) over threads (~44×); delayed TCP favored asyncio (~14× vs serial, ~2.4× vs threads w32); instant echo punished both concurrent arms; pure-Python CPU stayed GIL-bound (~0.92× with threads).
Evidence path on the lab box: lab-evidence/28-asyncio-vs-threads/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST on a shared Linux box (8 cores, kernel 6.12), Python 3.13.5. I ran 64 sleeps, delayed and instant 127.0.0.1 echo, mixed sleep plus slow TCP, and pure-Python CPU. Sleep: serial 1.301 s, ThreadPool 0.029 s, asyncio 0.022 s. Delayed TCP: asyncio 0.029 s and 2245 RPS; instant echo made serial win. Threads stayed flat on CPU (0.92×). No Docker, GPU, or API keys; affiliates 0.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026