ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 64

  1. Blog

shelve vs pickle Dict Store: Localhost Lab

A measured Linux localhost lab comparing shelve's keyed dbm store with a single pickle blob for full snapshots and random access.

Aditya Challa·30 September 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — n=8000 (p50)
  5. Scale sketch (write rec/s)
  6. Random access (200 keys @ n=8000)
  7. Reading it
  8. Durability tradeoff
  9. Why shelve write is slow here
  10. Snapshot vs store
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway

Intro — what this post promises

Persist a dict of records: shelve (dbm + pickle per key) vs one pickle.dump / load blob. This lab measures write/read rec/s and random-key access on Linux localhost.

It is not pickle vs JSON (lab 79). Here both arms use pickle encoding; the fork is keyed dbm store vs single blob.

Related links:

  • pickle vs json roundtrip localhost lab
  • weakref vs dict cache localhost lab
  • bytesio vs spooled tempfile localhost lab
  • secrets vs urandom localhost lab
  • tarfile vs zipfile localhost lab
  • scandir vs listdir localhost lab
  • futures as completed vs wait localhost lab
  • fnmatch vs re localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5, pickle.HIGHEST_PROTOCOL. Affiliates: 0. No Docker. Pickle is unsafe for untrusted data — local trusted stores only.

Verdict up front (n=8000): pickle write ~1523630 rec/s vs shelve write ~13436 (~113.0×); full read pickle ~1310138 vs shelve ~147828. 200 keyed gets: shelve ~1.97 ms vs pickle load-all-then-get ~7.54 ms (~3.8×).


Arms

ArmPattern
shelve.open writeone sh[k]=v per record + sync
pickle.dumpone blob file
shelve read alliterate keys / values
pickle read allload entire dict
random access200 keys via shelve vs load-all pickle

Lab topology

n in {500, 2000, 8000} records · 5 rounds · p50
metric: records/s ; random-access wall for 200 keys

Script: lab-evidence/94-shelve-vs-pickle-dict/results/run_lab.py.


Lead table — n=8000 (p50)

Armsrec/s
shelve write0.59513436
pickle write0.00531523630
shelve read all0.0541147828
pickle read all0.00611310138

Disk: shelve ~1568768 B vs pickle blob ~979971 B.


Scale sketch (write rec/s)

nshelve writepickle write
50040021815066
200091691691878
8000134361523630

Random access (200 keys @ n=8000)

Armp50 ms
shelve get keys1.97
pickle load entire dict then get7.54

Reading it

  • Full snapshot dump/load: pickle blob wins by a wide margin (one serialization pass).
  • Point lookups without loading everything: shelve wins — that is the durability/random-access product.
  • Shelve still pickles each value — untrusted input remains unsafe.
  • Prefer pickle for replace-the-world checkpoints; prefer shelve/SQLite when you need keyed updates.

Durability tradeoff

Shelve/dbm keeps a file you can open and fetch one key. A pickle blob is atomic only if you write-temp-and-rename (not measured here). Neither replaces a real database for concurrent writers.


Why shelve write is slow here

Each assignment pickles a value and updates the dbm map — fine for occasional writes, painful for bulk ingest of thousands of new keys. Bulk-load patterns usually build a dict then pickle once, or use SQLite executemany.


Snapshot vs store

Think of pickle as save game (replace the whole world) and shelve as key/value cabinet (open drawer id_042). Mixing them — rewriting a shelve by deleting all keys then reinserting — usually loses to dump-a-blob. Conversely, shipping a multi-megabyte pickle across the network to change one field is the wrong cabinet.


Pitfalls

  • Using shelve for hot full-table scans (pay per-key overhead).
  • Loading a huge pickle just to read one key.
  • Forgetting pickle trust boundaries.
  • Assuming shelve is cross-process safe without locking.

Reproduce

python3 lab-evidence/94-shelve-vs-pickle-dict/results/run_lab.py

Evidence: summary.json, bench_shelf_*, bench_blob_*.pkl.


Limits

One Linux box, default dbm backend. Not gdbm tuning, not SQLite. Pickle protocol = HIGHEST.


Takeaway

Pickle blob crushed full write/read (~1523630 vs ~13436 rec/s write ). Shelve won keyed access (~1.97 ms vs ~7.54 ms for 200 keys). Choose blob for snapshots; shelve when random access matters.

shelvepickledbmdictpythonlocalhostsre

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. n=8000: pickle write 1523630 rec/s vs shelve 13436; read pickle 1310138 vs shelve 147828. 200-key access: shelve 1.97ms vs pickle load-all 7.54ms. Not JSON lab 79. Affiliates: 0. Evidence: lab-evidence/94-shelve-vs-pickle-dict/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 06

    pickle vs json: Round-trip Lab

    Hands-on pickle vs json local round-trip lab: real throughput and size for modest dict/list payloads (trusted data), measured on Linux localhost (lab).

    30 Sept 2026

  • Plate 46

    setdefault vs defaultdict: Insert Lab

    Hands-on setdefault vs if-not-in vs defaultdict lab: real ops/s for list-append and counter default insert patterns, measured on Linux localhost (lab).

    Observability & SRE · 30 Sept 2026

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — n=8000 (p50)
  5. Scale sketch (write rec/s)
  6. Random access (200 keys @ n=8000)
  7. Reading it
  8. Durability tradeoff
  9. Why shelve write is slow here
  10. Snapshot vs store
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove