ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 96

  1. Blog

Queue vs deque Threaded Handoff: Localhost Lab

Aditya Challa·30 September 2026·4 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 200 000 messages (p50 Mmsgs/s)
  5. Scale sketch
  6. API vs speed
  7. Backpressure note
  8. Sentinel pattern
  9. Reading it
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway

Intro — what this post promises

Measure producer→consumer handoff across threads: queue.Queue, queue.SimpleQueue, and a collections.deque + Lock/Condition hand-roll. This lab reports messages/s on Linux localhost.

It is not the single-threaded deque-vs-list queue bake-off (lab 54) — that was same-thread push/pop. Here the cost is cross-thread synchronization.

Related links:

  • deque vs list queue localhost lab
  • groupby vs manual localhost lab
  • methodcaller vs getattr localhost lab
  • hmac compare digest localhost lab
  • argparse vs sys argv localhost lab
  • shelve vs pickle dict localhost lab
  • subprocess run vs popen localhost lab
  • futures as completed vs wait localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. No Docker. One producer + one consumer; sentinel None.

Verdict up front (200 000 msgs): SimpleQueue ~14.13 Mmsgs/s; deque+Lock batch64 ~3.09; deque notify-each ~1.58; Queue unbounded ~0.97; Queue maxsize=1024 ~0.88. Use Queue when you need blocking/task_done/join; SimpleQueue or a careful deque+Lock when both ends are yours and you want speed.


Arms

ArmPattern
Queue()classic unbounded
Queue(maxsize=1024)bounded backpressure
SimpleQueuesimple FIFO, no task_done
deque+Lock notify-eachCondition wake per put
deque+Lock batch64notify every 64 puts

Lab topology

n in {50k, 200k, 500k} ints · 1 producer · 1 consumer · 7 rounds · p50
metric: msgs/s = n / p50_s

Script: lab-evidence/100-queue-vs-deque-handoff/results/run_lab.py.


Lead table — 200 000 messages (p50 Mmsgs/s)

ArmMmsgs/s
SimpleQueue14.13
deque+Lock batch643.09
deque+Lock notify-each1.58
Queue unbounded0.97
Queue maxsize=10240.88

SimpleQueue leads by a wide margin on this box. Classic Queue pays for a richer API (mutex + conditions + unfinished-task tracking). Bounded maxsize adds a little more contention under a tight producer.


Scale sketch

nQueueSimpleQueuedeque batch64
50 k0.9815.933.22
200 k0.9714.133.09
500 k0.9614.53.04

Ordering holds across sizes — good sign the ranking is not a one-shot noise spike.


API vs speed

  • Queue: get/put timeouts, task_done/join, maxsize backpressure — pick when workers need join semantics.
  • SimpleQueue: fast FIFO for “just hand me items”; no unfinished-task counter.
  • deque+Lock: only when you control both ends and accept Condition bugs if you get wait/notify wrong. Batching notifies helps; notify-every-put approaches Queue cost.

Backpressure note

Queue(maxsize=1024) is slightly slower here because the producer sometimes blocks when the buffer fills. That is the point of bounded queues in production: protect memory at the cost of a few messages/s. Unbounded Queue and SimpleQueue can grow without limit under a slow consumer — fine for labs, risky for long-running services.


Sentinel pattern

Both ends agree on a poison-pill (None here). Without it, get() waits forever after the last item. For typed payloads, use a dedicated sentinel object, not a value that might appear in data.


Reading it

  • Prefer Queue for worker pools that join() on completion.
  • Prefer SimpleQueue for high-rate handoff without join bookkeeping.
  • Hand-rolled deque+Lock can beat classic Queue but rarely beats SimpleQueue here — and is easier to get wrong.
  • Lab 54’s single-thread deque speed does not transfer to threaded handoff.

Pitfalls

  • Using deque without a lock across threads (data races).
  • Forgetting a sentinel / poison pill and hanging forever.
  • Measuring only unbounded Queue then shipping bounded workers.
  • Confusing this with single-thread deque benches (lab 54).

Reproduce

python3 lab-evidence/100-queue-vs-deque-handoff/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box. Int payloads only. Two threads. No multi-consumer fan-out. GIL still serializes Python bytecode — this measures sync overhead, not parallel CPU.


Takeaway

At 200 k messages, SimpleQueue ~14.13 Mmsgs/s led; deque batch ~3.09; classic Queue ~0.97. Pick Queue for the API; pick SimpleQueue (or a careful deque) when handoff rate is the product metric.

pythonmultithreadingqueuedequeperformanceconcurrencysynchronization

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. 200k msgs: SimpleQueue 14.13 Mmsgs/s; deque batch64 3.09; Queue unbounded 0.97; Queue maxsize1024 0.88. Cross-thread (not lab 54). Affiliates: 0. Evidence: lab-evidence/100-queue-vs-deque-handoff/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 16

    ExitStack vs Nested with Resources: Localhost Lab

    Hands-on contextlib.ExitStack vs nested with and manual close: real cycles/s for N resources, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 57

    memoryview vs bytes Slice: Localhost Lab

    1 Oct 2026

  • Plate 82

    shlex.split vs str.split: Localhost Lab

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 200 000 messages (p50 Mmsgs/s)
  5. Scale sketch
  6. API vs speed
  7. Backpressure note
  8. Sentinel pattern
  9. Reading it
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove