ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 80

  1. Blog
  2. /Observability & SRE

Asyncio vs Threads: I/O Concurrency Lab with Real Numbers

Hands-on asyncio vs ThreadPool lab: sleep 58.7× vs 44×; delayed TCP 13.8× vs serial; fast-echo trap loses; CPU flat. Real localhost numbers, no Docker.

Aditya Challa·30 September 2026·6 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — concurrent sleep (the textbook win)
  5. Arm B — delayed TCP echo (5 ms think time)
  6. Arm C — instant echo (the overhead trap)
  7. Arm D — mixed sleep + slow TCP
  8. Arm E — pure-Python CPU (GIL control)
  9. When to pick which
  10. Pitfalls we hit (or avoided)
  11. Practical checklist
  12. Verdict

Intro — what this post promises

Asyncio and threads both overlap waiting work. They are not interchangeable: one multiplexes on a single event loop; the other pays thread stacks and the GIL tax when you accidentally do CPU.

This is a hands-on lab with measured numbers:

  1. Concurrent sleep — ideal wall ≈ one delay, not N×delay.
  2. Delayed localhost TCP echo — where concurrency actually wins.
  3. Instant echo — the overhead trap where serial wins.
  4. Mixed sleep + slow TCP on one loop vs a thread pool.
  5. Pure-Python CPU control — threads do not parallelize it.

Related links:

  • ProcessPool vs ThreadPool GIL localhost lab
  • Pipe vs tmpfile IPC localhost lab
  • epoll vs select FD_SETSIZE localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5 asyncio + ThreadPoolExecutor against a 127.0.0.1 echo (instant and 5 ms delayed). No Docker. No GPU. No API keys. Affiliates: 0.

Verdict up front: sleep asyncio ~59× / threads ~44× vs ideal serial; delayed TCP asyncio ~14× vs serial (~2.4× vs threads w32); instant echo concurrent arms lost; pure-Python CPU threads ~0.92×.


What we compared

ArmModelMetric
Aserial / threads / asyncio sleepwall vs ideal N×delay
Bdelayed TCP echo (5 ms server think)wall RPS + speedup
Cinstant TCP echo (overhead trap)RPS vs serial
Dmixed sleep + delayed TCPwall speedup
Epure-Python CPUthreads vs serial (GIL)

Related links:

  • asyncio — Concurrent code
  • concurrent.futures — ThreadPoolExecutor

Lab topology

Arm A: 64 × sleep(20 ms) — serial vs ThreadPool(64) vs asyncio.gather
Arm B/C: threaded echo on 127.0.0.1 — delay 5 ms vs 0 ms
         client: serial loop | ThreadPool | asyncio.open_connection gather
Arm D: 32×sleep(20 ms) + 32×slow TCP in one gather / one pool
Arm E: 8× pure-Python CPU loops — serial vs ThreadPool(8)

Arm A — concurrent sleep (the textbook win)

n=64, delay=20 ms, ideal serial = 1.280 s.

ModelWall sSpeedup vs ideal
serial1.301~1.0×
ThreadPool0.02944.4×
asyncio0.02258.7×

Both models collapse wall time toward one delay. Asyncio edged threads here: no per-task OS thread.


Arm B — delayed TCP echo (5 ms think time)

n=64 connect-per 64 B echo. Ideal serial wait ≈ 0.32 s (delay alone).

ModelWall sRPSvs serial
serial0.3931631.0×
threads w=80.05811076.8×
threads w=160.04215159.3×
threads w=320.0679545.9×
threads w=640.05112477.7×
asyncio0.029224513.8×

When each request waits on the peer, overlapping waits is the whole game. Asyncio beat the best thread pool by ~2.4× vs w32 on this box (thread-count sweet spot was w16, not “max workers = n”).

Related links:

  • Unix Domain Socket vs TCP localhost lab

Arm C — instant echo (the overhead trap)

Same shape, 0 ms server delay, n=200:

ModelRPSvs serial
serial26581.0×
threads w3220160.76×
asyncio13370.50×

Localhost connect+echo is so cheap that thread/task setup loses to a tight serial loop. Concurrency is not free — measure the wait you are overlapping.


Arm D — mixed sleep + slow TCP

32× sleep(20 ms) + 32× delayed TCP:

ModelWall svs serial
serial0.8351.0×
ThreadPool0.02928.6×
asyncio0.02139.1×

One event loop (or one pool) overlaps both wait types. This is the “many sockets + timers” shape asyncio is for.


Arm E — pure-Python CPU (GIL control)

8 tasks of tight pure-Python arithmetic:

ModelWall svs serial
serial0.0471.0×
ThreadPool(8)0.0510.92×

Threads did not speed this up. For CPU-bound pure Python, use processes (see the GIL pool lab) — not more threads, and not await on hot loops.

Related links:

  • ProcessPool vs ThreadPool GIL localhost lab

When to pick which

SituationPrefer
Many waits (sockets, sleeps, timers)asyncio (or threads if the ecosystem is sync-only)
Blocking sync library you cannot rewriteThreadPoolExecutor off the loop
Pure-Python / GIL-bound CPUProcessPool (not this post’s winners)
Microsecond localhost RPC with no waitMeasure — serial may win (Arm C)

Pitfalls we hit (or avoided)

  1. Benchmarking only instant localhost echo — Arm C lies about I/O concurrency value.
  2. Assuming more thread workers always help — w16 beat w32/w64 on Arm B.
  3. Putting CPU work on the event loop — serializes everything behind it.
  4. Treating this as uvloop / Trio / gevent — we measured stdlib asyncio + threads only.
  5. Confusing with ProcessPool CPU scaling — different problem; different lab.

Practical checklist

  • Name the wait you are overlapping (delay ms, RTT, disk). If wait ≈ 0, concurrency may lose.
  • Report wall + RPS + worker/concurrency + serial baseline together.
  • Cap thread workers; sweep (8/16/32) instead of workers=n_tasks by habit.
  • Keep CPU off the loop; hand it to processes or native extensions that release the GIL.
  • Re-run under your real client library — wrappers add their own tax.

Verdict

On this box: concurrent sleep favored asyncio (~59×) over threads (~44×); delayed TCP favored asyncio (~14× vs serial, ~2.4× vs threads w32); instant echo punished both concurrent arms; pure-Python CPU stayed GIL-bound (~0.92× with threads).

Evidence path on the lab box: lab-evidence/28-asyncio-vs-threads/results/. Affiliates: 0.

asyncio vs threadspython asyncio concurrencythreadpoolexecutor i/oevent loop vs threadslocalhost labsreconcurrent i/ogil i/o bound

Lab evidence

What I found running this

Lab 1 Oct 2026 IST on a shared Linux box (8 cores, kernel 6.12), Python 3.13.5. I ran 64 sleeps, delayed and instant 127.0.0.1 echo, mixed sleep plus slow TCP, and pure-Python CPU. Sleep: serial 1.301 s, ThreadPool 0.029 s, asyncio 0.022 s. Delayed TCP: asyncio 0.029 s and 2245 RPS; instant echo made serial win. Threads stayed flat on CPU (0.92×). No Docker, GPU, or API keys; affiliates 0.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — concurrent sleep (the textbook win)
  5. Arm B — delayed TCP echo (5 ms think time)
  6. Arm C — instant echo (the overhead trap)
  7. Arm D — mixed sleep + slow TCP
  8. Arm E — pure-Python CPU (GIL control)
  9. When to pick which
  10. Pitfalls we hit (or avoided)
  11. Practical checklist
  12. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove